AI Dictionary › Fondamenti AI

Evaluation metrics

Metriche di valutazione

Evaluation metrics are the numeric measures used to quantify how well a machine learning model performs its task, allowing different models to be compared, the best one to be chosen, and a decision made on whether the achieved performance is sufficient for real-world use. Without a shared metric, claiming a model works well is a vague, unverifiable statement.

Definition

What it is

Evaluation metrics are the numeric measures used to quantify how well a machine learning model performs its task, allowing different models to be compared, the best one to be chosen, and a decision made on whether the achieved performance is sufficient for real-world use. Without a shared metric, claiming a model works well is a vague, unverifiable statement.

How it works

The choice of metric depends on the type of problem. For regression, measures like mean squared error or mean absolute error are used, quantifying how far predictions deviate from actual values. For classification, accuracy, precision, recall, F1-score and area under the ROC curve are used, each sensitive to a different aspect of the errors made. There is no universally best metric: the choice depends on what truly matters for the specific problem, for example whether a false positive or a false negative is worse.

Applications

In applied AI, metrics guide every technical decision: which model to put into production, when an update actually represents an improvement, when a system has degraded over time and needs retraining. They are also a communication tool between technical teams and business decision-makers, who rarely understand algorithmic details but can grasp that precision rose from 80 to 90 percent.

History & etymology

The systematic use of quantitative measures to evaluate the quality of statistical predictions belongs to early twentieth-century statistics, but the formalization of the metrics standard in machine learning today took shape starting in the 1960s and 70s, in fields like pattern recognition and information theory, before spreading widely with the discipline's growth in the 1990s and 2000s.

Definizione (italiano)

Le metriche di valutazione sono le misure numeriche usate per quantificare quanto bene un modello di machine learning svolge il proprio compito, permettendo di confrontare modelli diversi, scegliere il migliore e decidere se le prestazioni raggiunte sono sufficienti per un uso reale. Senza una metrica condivisa, dire che un modello funziona bene è un'affermazione vaga e non verificabile.

La scelta della metrica dipende dal tipo di problema. Per la regressione si usano misure come l'errore quadratico medio o l'errore assoluto medio, che quantificano quanto le previsioni si discostano dai valori reali. Per la classificazione si usano accuratezza, precisione, richiamo, F1-score e area sotto la curva ROC, ciascuna sensibile a un aspetto diverso degli errori commessi. Non esiste una metrica universalmente migliore: la scelta dipende da cosa conta davvero per il problema specifico, per esempio se è peggio un falso positivo o un falso negativo.

Nell'AI applicata le metriche guidano ogni decisione tecnica: quale modello mettere in produzione, quando un aggiornamento rappresenta davvero un miglioramento, quando un sistema si è degradato nel tempo e va riaddestrato. Sono anche uno strumento di comunicazione tra i team tecnici e chi prende decisioni di business, che raramente comprende i dettagli dell'algoritmo ma può capire se la precisione è salita dall'80 al 90 per cento.

L'uso sistematico di misure quantitative per valutare la qualità di previsioni statistiche appartiene alla statistica del primo Novecento, ma la formalizzazione delle metriche oggi standard nel machine learning si consolida a partire dagli anni '60 e '70, in ambiti come il riconoscimento di pattern e la teoria dell'informazione, per poi diffondersi ampiamente con la crescita della disciplina negli anni '90 e 2000.

Related terms

More in Fondamenti AI

Put it into practice

From our network

Magellano GPS — Fleet Tracking Made Simple

Real-time GPS tracking, remote engine lock, fuel and CO₂ reporting for your fleet.

Visit magellanogps.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.