AI Dictionary › Fondamenti AI
Task di valutazione (Evaluation Task)
An evaluation task is a specific job used to measure a precise capability of an AI model, such as answering questions, summarizing a text, translating, or solving math problems. It is the basic unit from which benchmarks are built: a benchmark is typically made up of many similar tasks grouped by area of competence. Defining the task well, including instructions and expected response format, is essential for a reliable measurement.
Each task includes an input, a clear request, and a criterion for judging whether the model's response is correct or of good quality. The criterion can be automated, such as comparison with a reference answer, or require human judgment when correctness is more nuanced. Results on individual tasks are often aggregated into an overall score for a competence area.
Evaluation tasks are used to test targeted capabilities before a model's release, to diagnose which areas a system is weaker in, and to build custom test suites for specific business needs. They are also the tool used to check for regressions after a model update.
The concept has roots in the tradition of evaluation in computational linguistics and machine learning, where standardized tasks were already used before the spread of large language models to objectively compare different systems.
Every Grace scenario is built as an evaluation task: it defines an input, a clear request and correctness criteria on which the user's score is calculated.
From our network
Magellano GPS: Fleet Tracking Made Simple
Real-time GPS tracking, remote engine lock, fuel and CO₂ reporting for your fleet.
Visit magellanogps.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.