AI Dictionary › Prompting
Filtro di Sicurezza
A safety filter is a software component, separate from the language model itself, that inspects the input and output of an AI system for prohibited, dangerous, or policy-violating content. It can block a request before it reaches the model, or intercept and modify a response before it is shown to the user.
Technically, safety filters rely on dedicated classifiers, often smaller and faster models trained to recognize risk categories (violence, sexual content, self-harm, dangerous instructions), or on word lists and regular expressions for simpler cases. Many systems combine multiple layers: a lightweight, fast filter on input, a deeper check on generated output, and sometimes a final verification before delivery to the user.
In modern AI, safety filters are the layer that separates a capable but potentially wrong or manipulable model from a product usable in production: the commercial APIs of major providers apply them by default and often make them configurable for sectors with different needs (healthcare, adult content, educational applications).
The concept derives from user-generated content moderation systems, already widespread on social media before the era of large language models, and has been adapted and made more sophisticated with the arrival of generative AI systems capable of producing text, images, and code on demand.
From our network
INDACO TMS: Transport Management for European Logistics
Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.
Visit indacotms.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.