AI Dictionary › Fondamenti AI
Task di valutazione (Evaluation Task)
An evaluation task is a specific job used to measure a precise capability of an AI model, such as answering questions, summarizing a text, translating, or solving math problems. It is the basic unit from which benchmarks are built: a benchmark is typically made up of many similar tasks grouped by area of competence. Defining the task well, including instructions and expected response format, is essential for a reliable measurement.
An evaluation task is a specific job used to measure a precise capability of an AI model, such as answering questions, summarizing a text, translating, or solving math problems. It is the basic unit from which benchmarks are built: a benchmark is typically made up of many similar tasks grouped by area of competence. Defining the task well, including instructions and expected response format, is essential for a reliable measurement.
Each task includes an input, a clear request, and a criterion for judging whether the model's response is correct or of good quality. The criterion can be automated, such as comparison with a reference answer, or require human judgment when correctness is more nuanced. Results on individual tasks are often aggregated into an overall score for a competence area.
Evaluation tasks are used to test targeted capabilities before a model's release, to diagnose which areas a system is weaker in, and to build custom test suites for specific business needs. They are also the tool used to check for regressions after a model update.
The concept has roots in the tradition of evaluation in computational linguistics and machine learning, where standardized tasks were already used before the spread of large language models to objectively compare different systems.
Un task di valutazione è un compito specifico usato per misurare una capacità precisa di un modello AI, come rispondere a domande, riassumere un testo, tradurre o risolvere problemi matematici. È l'unità di base con cui si costruiscono i benchmark: ogni benchmark è tipicamente composto da molti task simili raggruppati per area di competenza. Definire bene il task, incluse le istruzioni e il formato atteso della risposta, è essenziale per ottenere una misura affidabile.
Ogni task include un input, una richiesta chiara e un criterio per stabilire se la risposta del modello è corretta o di buona qualità. Il criterio può essere automatico, come il confronto con una risposta di riferimento, oppure richiedere una valutazione umana quando la correttezza è più sfumata. I risultati sui singoli task vengono spesso aggregati per ottenere un punteggio complessivo su un'area di competenza.
I task di valutazione vengono usati per testare capacità mirate prima del rilascio di un modello, per diagnosticare in quali aree un sistema è più debole e per costruire suite di test personalizzate su esigenze aziendali specifiche. Sono anche lo strumento con cui si verificano eventuali regressioni dopo un aggiornamento del modello.
Il concetto affonda le radici nella tradizione della valutazione in linguistica computazionale e apprendimento automatico, dove compiti standardizzati venivano usati già prima della diffusione dei grandi modelli linguistici per confrontare sistemi diversi in modo oggettivo.
From our network
Magellano GPS — Fleet Tracking Made Simple
Real-time GPS tracking, remote engine lock, fuel and CO₂ reporting for your fleet.
Visit magellanogps.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.