AI Dictionary › Fondamenti AI

Human Evaluation

Valutazione umana (Human Evaluation)

Human evaluation is the method by which real people judge the quality of an AI system's outputs, complementing or replacing automated metrics. It is especially useful for aspects that are hard to quantify with a single number, such as the naturalness of a text, the logical coherence of a response, or how useful it feels to the recipient. It is often used when automated benchmarks fail to capture important nuances.

Definition

What it is

Human evaluation is the method by which real people judge the quality of an AI system's outputs, complementing or replacing automated metrics. It is especially useful for aspects that are hard to quantify with a single number, such as the naturalness of a text, the logical coherence of a response, or how useful it feels to the recipient. It is often used when automated benchmarks fail to capture important nuances.

How it works

In practice, human raters read or compare different outputs and assign a judgment according to predefined criteria, such as a numeric scale or a pairwise comparison between two responses. To reduce subjectivity, multiple raters assess the same output and inter-rater agreement is calculated. Results are then aggregated into overall scores or preference rankings.

Applications

It is widely used to validate language models before release, to train systems through human feedback, and to verify that AI responses are appropriate in sensitive contexts such as health or professional advice. Many evaluation platforms pair it with automated tests for a fuller picture.

History & etymology

It has always accompanied research on AI and natural language, but became central with the spread of conversational models, when the quality perceived by real people took on decisive weight in judging a system.

Definizione (italiano)

La valutazione umana è il metodo con cui persone reali giudicano la qualità degli output di un sistema AI, integrando o sostituendo le metriche automatiche. È particolarmente utile per aspetti difficili da quantificare con un numero, come la naturalezza di un testo, la coerenza logica di una risposta o la percezione di utilità da parte di chi la riceve. Viene spesso usata quando i benchmark automatici non catturano sfumature importanti.

In pratica, valutatori umani leggono o confrontano output diversi e assegnano un giudizio secondo criteri predefiniti, ad esempio una scala numerica o un confronto a coppie tra due risposte. Per ridurre la soggettività si usano più valutatori sullo stesso output e si calcola l'accordo tra loro. I risultati vengono poi aggregati in punteggi complessivi o in classifiche di preferenza.

È ampiamente usata per validare modelli linguistici prima del rilascio, per addestrare sistemi tramite feedback umano e per verificare che le risposte AI siano appropriate in contesti sensibili come salute o consulenza professionale. Molte piattaforme di valutazione la affiancano ai test automatici per avere un quadro più completo.

È una pratica che affianca da sempre la ricerca su AI e linguaggio naturale, ma è diventata centrale con la diffusione dei modelli conversazionali, quando la qualità percepita da persone reali ha assunto un peso decisivo nel giudicare un sistema.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.