AI Dictionary › Prompting
Rilevamento della Tossicità
Toxicity detection is the specific task, within AI content moderation, of identifying offensive, aggressive, discriminatory, or harassing language in a text, typically assigning it a toxicity score on a continuous scale rather than a simple binary label. This granularity allows thresholds to be calibrated differently depending on context: an adult community may tolerate harsher language than an educational application for minors.
Toxicity detection systems are trained on datasets of comments and text annotated by human raters according to guidelines defining what counts as toxic, often distinguishing sub-categories such as threats, insults, obscene language, or identity-based attacks. The model learns to recognize linguistic patterns associated with these categories, including variants that evade simpler filters through intentional typos or character substitutions.
It is widely used to moderate comments on online platforms, filter the output of public chatbots, and assess the quality of data used to train new language models, since toxic text in the training set tends to produce undesirable behavior in the resulting model. Several open-source tools and commercial APIs offer ready-to-use toxicity scores that can be integrated into broader moderation pipelines.
The term and the first dedicated tools date back to the late 2010s, when tech companies handling large volumes of user-generated comments began investing in automated classifiers to complement human moderation, which was no longer sufficient on its own to cover the volumes involved.
From our network
AGORÀ Intelligence: Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.