AI Dictionary › Modelli AI

LLM-as-Judge

LLM come giudice (LLM-as-Judge)

LLM-as-judge is an evaluation technique in which a language model is used to judge the quality of responses generated by another model, in place of or alongside human evaluators. The judge model receives the original question, the response to be evaluated and explicit criteria, and returns a score or a comparison between alternatives. It has become popular because it allows large volumes of output to be evaluated without the cost and time of human review.

Definition

What it is

LLM-as-judge is an evaluation technique in which a language model is used to judge the quality of responses generated by another model, in place of or alongside human evaluators. The judge model receives the original question, the response to be evaluated and explicit criteria, and returns a score or a comparison between alternatives. It has become popular because it allows large volumes of output to be evaluated without the cost and time of human review.

How it works

In practice, the judge model is given a structured prompt with evaluation instructions, often accompanied by reference examples. The model analyzes the response according to criteria such as correctness, clarity or adherence to instructions, and produces a reasoned judgment. To increase reliability, multiple judge models can be used, or their outcomes can be compared against a sample evaluated by people.

Applications

It is used to quickly test new model versions, to compare outputs of different systems at scale, and to build continuous evaluation pipelines during AI product development. It should still be used carefully, since the judge can inherit the same limitations or biases as the models it evaluates.

History & etymology

It is a practice that emerged as large language models matured, when their text comprehension capabilities became reliable enough to be employed in evaluation tasks as well.

Definizione (italiano)

LLM-as-judge è una tecnica di valutazione in cui un modello linguistico viene usato per giudicare la qualità delle risposte generate da un altro modello, al posto o in aggiunta a valutatori umani. Il modello giudice riceve la domanda originale, la risposta da valutare e criteri espliciti, e restituisce un punteggio o un confronto tra alternative. È diventata popolare perché permette di valutare grandi volumi di output senza il costo e i tempi della revisione umana.

In pratica si fornisce al modello giudice un prompt strutturato con le istruzioni di valutazione, spesso accompagnato da esempi di riferimento. Il modello analizza la risposta secondo criteri come correttezza, chiarezza o aderenza alle istruzioni e produce un giudizio motivato. Per aumentare l'affidabilità si possono usare più modelli giudice o confrontare i loro esiti con un campione valutato da persone.

È utilizzata per testare rapidamente nuove versioni di un modello, per confrontare output di sistemi diversi su larga scala e per creare pipeline di valutazione continua durante lo sviluppo di un prodotto AI. Va comunque usata con attenzione, perché il giudice può ereditare gli stessi limiti o pregiudizi dei modelli che valuta.

È una pratica emersa con la maturazione dei modelli linguistici di grandi dimensioni, quando le loro capacità di comprensione del testo sono diventate sufficientemente affidabili da essere impiegate anche in compiti di valutazione.

Related terms

More in Modelli AI

Put it into practice

From our network

Kaimaki Web — Websites That Win Customers

Custom websites, web apps and digital marketing for growing businesses.

Visit kaimakiweb.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.