AI Dictionary › Fondamenti AI
Addestramento al Rifiuto
Refusal training is the stage of the alignment process in which a language model is explicitly taught to recognize harmful, dangerous, or policy-violating requests and to respond with a refusal instead of attempting to fulfill them. It is one of the most visible components of a model's safety behavior: when a chatbot answers "I can't help with that", that behavior is typically the direct result of this kind of training.
Refusal training is the stage of the alignment process in which a language model is explicitly taught to recognize harmful, dangerous, or policy-violating requests and to respond with a refusal instead of attempting to fulfill them. It is one of the most visible components of a model's safety behavior: when a chatbot answers "I can't help with that", that behavior is typically the direct result of this kind of training.
Technically it is implemented through supervised learning on pairs of harmful requests and appropriate refusal responses, often combined with reinforcement learning from human feedback in which raters reward well-calibrated refusals and penalize both excessive refusal of harmless requests and compliance with genuinely dangerous ones. Striking this balance is technically difficult: a model that refuses too much becomes unhelpful, while one that refuses too little exposes concrete risks.
It is continuously tested by red-teaming teams, who look for phrasings able to bypass refusal through role-play, hypothetical scenarios, or requests broken into seemingly innocuous steps. AI labs periodically update refusal training as new evasion techniques emerge, in an iterative process that never truly concludes.
The concept emerged with the spread of large conversational language models starting in 2022, when it became clear that a model capable of generating text on any topic needed specific, separate training to learn how to say no appropriately.
L'addestramento al rifiuto è la fase del processo di allineamento in cui un modello linguistico viene esplicitamente insegnato a riconoscere richieste dannose, pericolose o contrarie alle policy d'uso e a rispondere con un rifiuto anziché tentare di soddisfarle. È una delle componenti più visibili del comportamento di sicurezza di un modello: quando un chatbot risponde "non posso aiutarti con questo", quel comportamento è tipicamente il risultato diretto di questo tipo di addestramento.
Tecnicamente si realizza tramite apprendimento supervisionato su coppie di richieste dannose e risposte di rifiuto appropriate, spesso combinato con reinforcement learning from human feedback in cui i valutatori premiano rifiuti ben calibrati e penalizzano sia il rifiuto eccessivo di richieste innocue sia la compliance con richieste effettivamente pericolose. Trovare questo equilibrio è tecnicamente difficile: un modello che rifiuta troppo diventa poco utile, mentre uno che rifiuta troppo poco espone a rischi concreti.
È oggetto continuo di test da parte dei team di red teaming, che cercano formulazioni capaci di aggirare il rifiuto tramite giochi di ruolo, scenari ipotetici o richieste frammentate in passaggi apparentemente innocui. I laboratori AI aggiornano periodicamente l'addestramento al rifiuto man mano che emergono nuove tecniche di elusione, in un processo iterativo che non si conclude mai definitivamente.
Il concetto è emerso con la diffusione dei grandi modelli linguistici conversazionali a partire dal 2022, quando è diventato evidente che un modello capace di generare testo su qualsiasi argomento richiedeva un addestramento specifico e separato per imparare a dire no in modo appropriato.
From our network
HSE Genius — AI for Safety Data Sheets
Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.
Visit hsegenius.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.