AI Dictionary › Fondamenti AI

Safety Benchmark

Benchmark di Sicurezza

A safety benchmark is a standardized set of tests used to measure, in a comparable way, how well an AI model resists harmful requests, manipulation attempts, and risky situations, much like a traditional benchmark measures capabilities such as mathematical reasoning or text comprehension. Unlike performance benchmarks, which reward a capable and useful model, safety benchmarks reward a model that knows when to say no without becoming needlessly restrictive.

Definition

What it is

A safety benchmark is a standardized set of tests used to measure, in a comparable way, how well an AI model resists harmful requests, manipulation attempts, and risky situations, much like a traditional benchmark measures capabilities such as mathematical reasoning or text comprehension. Unlike performance benchmarks, which reward a capable and useful model, safety benchmarks reward a model that knows when to say no without becoming needlessly restrictive.

How it works

Technically, a safety benchmark consists of large collections of prompts designed to represent different risk categories — explicitly harmful requests, known jailbreak attempts, ambiguous requests on the border between lawful and unlawful, and harmless questions an overly cautious model might mistakenly refuse. The model is made to respond to each prompt, and the responses are evaluated, often with the help of other AI models used as automated judges, against predefined safety and helpfulness criteria.

Applications

It is a central tool for comparing different models before adopting them in production, for tracking a model's improvements from one version to the next, and for meeting regulatory requirements that call for a documented risk assessment before release. Several safety benchmarks are now public and widely cited in the technical reports of major AI labs, enabling a transparent comparison between competing providers.

History & etymology

These tools developed as a distinct category starting in 2022-2023, when the mass adoption of public chatbots made it necessary to have a systematic, repeatable way of measuring a model's safety, instead of relying solely on manual case-by-case testing.

Definizione (italiano)

Un benchmark di sicurezza è un set standardizzato di test usato per misurare in modo comparabile quanto un modello AI resista a richieste dannose, tentativi di manipolazione e situazioni rischiose, allo stesso modo in cui un benchmark tradizionale misura capacità come ragionamento matematico o comprensione del testo. A differenza dei benchmark di performance, che premiano un modello capace e utile, i benchmark di sicurezza premiano un modello che sa dire no al momento giusto senza diventare inutilmente restrittivo.

Tecnicamente un benchmark di sicurezza è composto da grandi collezioni di prompt studiati per rappresentare categorie di rischio diverse — richieste esplicitamente dannose, tentativi di jailbreak noti, richieste ambigue al confine tra lecito e illecito, domande innocue che un modello troppo prudente potrebbe rifiutare per errore. Il modello viene fatto rispondere a ciascun prompt e le risposte vengono valutate, spesso con l'aiuto di altri modelli AI usati come giudici automatici, secondo criteri predefiniti di sicurezza e utilità.

È uno strumento centrale per confrontare modelli diversi prima di adottarli in produzione, per tracciare i miglioramenti di un modello tra una versione e l'altra, e per soddisfare requisiti normativi che richiedono una valutazione documentata dei rischi prima del rilascio. Diversi benchmark di sicurezza sono oggi pubblici e ampiamente citati nei report tecnici dei principali laboratori AI, permettendo un confronto trasparente tra fornitori concorrenti.

Questi strumenti si sono sviluppati come categoria distinta a partire dal 2022-2023, quando la diffusione di massa dei chatbot pubblici ha reso necessario un modo sistematico e ripetibile di misurare la sicurezza di un modello, invece di affidarsi solo a test manuali caso per caso.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.