AI Dictionary › Fondamenti AI

Evaluation Task

Task di valutazione (Evaluation Task)

An evaluation task is a specific job used to measure a precise capability of an AI model, such as answering questions, summarizing a text, translating, or solving math problems. It is the basic unit from which benchmarks are built: a benchmark is typically made up of many similar tasks grouped by area of competence. Defining the task well, including instructions and expected response format, is essential for a reliable measurement.

Definition

How it works

Each task includes an input, a clear request, and a criterion for judging whether the model's response is correct or of good quality. The criterion can be automated, such as comparison with a reference answer, or require human judgment when correctness is more nuanced. Results on individual tasks are often aggregated into an overall score for a competence area.

Applications

Evaluation tasks are used to test targeted capabilities before a model's release, to diagnose which areas a system is weaker in, and to build custom test suites for specific business needs. They are also the tool used to check for regressions after a model update.

History & etymology

The concept has roots in the tradition of evaluation in computational linguistics and machine learning, where standardized tasks were already used before the spread of large language models to objectively compare different systems.

How it's used in Grace

Every Grace scenario is built as an evaluation task: it defines an input, a clear request and correctness criteria on which the user's score is calculated.

Related terms

More in Fondamenti AI

Put it into practice

From our network

Magellano GPS: Fleet Tracking Made Simple

Real-time GPS tracking, remote engine lock, fuel and CO₂ reporting for your fleet.

Visit magellanogps.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.