AI Dictionary › Fondamenti AI

Robustness Testing

Test di robustezza (Robustness Testing)

Robustness testing checks how well an AI model maintains reliable performance when the input changes unexpectedly, for example through typos, unusual phrasing or small variations from the training data. A robust model keeps giving sensible responses even under non-ideal conditions, while a fragile one can change behavior drastically in the face of small perturbations. It is an aspect of evaluation distinct from plain accuracy on standard cases.

Definition

What it is

Robustness testing checks how well an AI model maintains reliable performance when the input changes unexpectedly, for example through typos, unusual phrasing or small variations from the training data. A robust model keeps giving sensible responses even under non-ideal conditions, while a fragile one can change behavior drastically in the face of small perturbations. It is an aspect of evaluation distinct from plain accuracy on standard cases.

How it works

To perform it, variants of the original input are generated, such as synonyms, rephrasings or artificial noise, and it is observed whether the model's response stays consistent and correct. Performance on the original input is then compared with performance on the variants, to quantify how sensitive the model is to change. Significant performance drops flag areas needing improvement.

Applications

It is used to validate models before production release, especially in contexts where real inputs are unpredictable, such as conversational assistants or decision-support systems. It also helps identify vulnerabilities exploitable by malicious inputs deliberately designed to fool the model.

History & etymology

It is a concept inherited from software engineering and control systems, where robustness has long referred to a system's ability to function correctly even under unforeseen or adverse conditions.

Definizione (italiano)

Il test di robustezza verifica quanto un modello AI mantenga prestazioni affidabili quando l'input cambia in modo inatteso, ad esempio con errori di battitura, formulazioni insolite o piccole variazioni rispetto ai dati di addestramento. Un modello robusto continua a fornire risposte sensate anche in condizioni non ideali, mentre uno fragile può cambiare drasticamente comportamento davanti a piccole perturbazioni. È un aspetto della valutazione distinto dalla semplice accuratezza su casi standard.

Per eseguirlo si generano varianti dell'input originale, come sinonimi, riformulazioni o rumore artificiale, e si osserva se la risposta del modello resta coerente e corretta. Si confrontano poi le prestazioni sull'input originale con quelle sulle varianti, per quantificare quanto il modello sia sensibile ai cambiamenti. Cali significativi di prestazione segnalano aree da migliorare.

È usato per validare modelli prima del rilascio in produzione, specialmente in contesti dove gli input reali sono imprevedibili, come assistenti conversazionali o sistemi di supporto decisionale. Aiuta anche a individuare vulnerabilità sfruttabili da input malevoli progettati apposta per ingannare il modello.

È un concetto ereditato dall'ingegneria del software e dai sistemi di controllo, dove la robustezza indica da tempo la capacità di un sistema di funzionare correttamente anche in condizioni impreviste o avverse.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.