AI Dictionary › Fondamenti AI

Model Evaluation

Valutazione del modello (Model Evaluation)

Model evaluation is the process of measuring how well an AI system performs the tasks it was designed for. It is not a single test but a set of quantitative and qualitative methods covering accuracy, robustness, safety and consistency of responses. It is used to compare different models, decide whether a model is ready for production, and spot weaknesses before they become real problems.

Definition

What it is

Model evaluation is the process of measuring how well an AI system performs the tasks it was designed for. It is not a single test but a set of quantitative and qualitative methods covering accuracy, robustness, safety and consistency of responses. It is used to compare different models, decide whether a model is ready for production, and spot weaknesses before they become real problems.

How it works

In practice, automated benchmarks, task-specific tests and human review of outputs are combined. The resulting scores are compared against reference thresholds or against other models' performance, giving a comparable picture rather than an isolated number. Evaluation is often repeated over time, since a model can degrade or improve with subsequent updates.

Applications

Companies use model evaluation to choose which AI to adopt, to certify product quality before release, and to monitor performance after deployment. It is also central in research, where new models are systematically compared against existing ones.

History & etymology

The term spread alongside the growth of large language models, when it became clear that performance could not be taken for granted and a shared vocabulary was needed to describe and compare it.

Definizione (italiano)

La valutazione del modello è il processo con cui si misura quanto bene un sistema AI svolge i compiti per cui è stato progettato. Non è un singolo test ma un insieme di metodi, quantitativi e qualitativi, che coprono accuratezza, robustezza, sicurezza e coerenza delle risposte. Serve a confrontare modelli diversi, a decidere se un modello è pronto per la produzione e a individuare punti deboli prima che diventino problemi reali.

In pratica si combinano benchmark automatici, test mirati su compiti specifici e revisione umana dei risultati. I punteggi ottenuti vengono confrontati con soglie di riferimento o con le prestazioni di altri modelli, così da avere un quadro comparabile e non solo un numero isolato. Spesso la valutazione viene ripetuta nel tempo, perché un modello può degradare o migliorare con aggiornamenti successivi.

Le aziende usano la valutazione del modello per scegliere quale AI adottare, per certificare la qualità di un prodotto prima del rilascio e per monitorare le prestazioni dopo la messa in produzione. È un passaggio centrale anche nella ricerca, dove nuovi modelli vengono confrontati sistematicamente con quelli esistenti.

Il termine si è diffuso insieme alla crescita dei modelli linguistici di grandi dimensioni, quando è diventato chiaro che le prestazioni non potevano essere date per scontate e serviva un vocabolario condiviso per descriverle e confrontarle.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.