AI Dictionary › AI Fundamentals
Addestramento al Rifiuto
Refusal training is the stage of the alignment process in which a language model is explicitly taught to recognize harmful, dangerous, or policy-violating requests and to respond with a refusal instead of attempting to fulfill them. It is one of the most visible components of a model's safety behavior: when a chatbot answers "I can't help with that", that behavior is typically the direct result of this kind of training.
Technically it is implemented through supervised learning on pairs of harmful requests and appropriate refusal responses, often combined with reinforcement learning from human feedback in which raters reward well-calibrated refusals and penalize both excessive refusal of harmless requests and compliance with genuinely dangerous ones. Striking this balance is technically difficult: a model that refuses too much becomes unhelpful, while one that refuses too little exposes concrete risks.
It is continuously tested by red-teaming teams, who look for phrasings able to bypass refusal through role-play, hypothetical scenarios, or requests broken into seemingly innocuous steps. AI labs periodically update refusal training as new evasion techniques emerge, in an iterative process that never truly concludes.
The concept emerged with the spread of large conversational language models starting in 2022, when it became clear that a model capable of generating text on any topic needed specific, separate training to learn how to say no appropriately.
From our network
HSE Genius: AI for Safety Data Sheets
Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.
Visit hsegenius.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.