AI Dictionary › Prompting
Classificatore di Contenuti Dannosi
A harmful content classifier is a model, often smaller and more specialized than the main generative model, trained to label text or images according to predefined risk categories: violence, hate speech, sexual content, self-harm, instructions for illegal activities. Its job is not to generate content but to judge it, returning a score or label that other system components use to decide whether to block, flag, or allow it through.
These classifiers are trained with supervised learning on large datasets of examples labeled by people as harmful or safe, often enriched with data augmentation techniques to cover linguistic variants, sarcasm, code words, and implicit references. Many providers offer these classifiers as a separate API from the generation model, so anyone building an application can integrate them without training one from scratch.
They are the technical component behind many moderation functions visible to the end user: a chatbot refusing to answer certain questions, a generated image being blocked before it is shown, a post on a social platform being automatically flagged. They are also used internally by AI labs to filter training data and reduce the presence of toxic content in models from the pre-training stage onward.
The concept originates in social media content moderation, well before the spread of large language models, and has been adapted to generative systems as these began producing content at scale.
From our network
AGORÀ Intelligence: Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.