AI Dictionary › AI Fundamentals

Model Evaluation

Valutazione del modello (Model Evaluation)

Model evaluation is the process of measuring how well an AI system performs the tasks it was designed for. It is not a single test but a set of quantitative and qualitative methods covering accuracy, robustness, safety and consistency of responses. It is used to compare different models, decide whether a model is ready for production, and spot weaknesses before they become real problems.

Definition

How it works

In practice, automated benchmarks, task-specific tests and human review of outputs are combined. The resulting scores are compared against reference thresholds or against other models' performance, giving a comparable picture rather than an isolated number. Evaluation is often repeated over time, since a model can degrade or improve with subsequent updates.

Applications

Companies use model evaluation to choose which AI to adopt, to certify product quality before release, and to monitor performance after deployment. It is also central in research, where new models are systematically compared against existing ones.

History & etymology

The term spread alongside the growth of large language models, when it became clear that performance could not be taken for granted and a shared vocabulary was needed to describe and compare it.

How it's used in Grace

Grace applies model evaluation to its own scenarios: every answer users generate is measured against comparable criteria, so scores stay consistent regardless of the AI model chosen.

Related terms

More in AI Fundamentals

Put it into practice

From our network

AGORÀ Intelligence: Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.