AI Dictionary › Modelli AI
LLM come giudice (LLM-as-Judge)
LLM-as-judge is an evaluation technique in which a language model is used to judge the quality of responses generated by another model, in place of or alongside human evaluators. The judge model receives the original question, the response to be evaluated and explicit criteria, and returns a score or a comparison between alternatives. It has become popular because it allows large volumes of output to be evaluated without the cost and time of human review.
LLM-as-judge is an evaluation technique in which a language model is used to judge the quality of responses generated by another model, in place of or alongside human evaluators. The judge model receives the original question, the response to be evaluated and explicit criteria, and returns a score or a comparison between alternatives. It has become popular because it allows large volumes of output to be evaluated without the cost and time of human review.
In practice, the judge model is given a structured prompt with evaluation instructions, often accompanied by reference examples. The model analyzes the response according to criteria such as correctness, clarity or adherence to instructions, and produces a reasoned judgment. To increase reliability, multiple judge models can be used, or their outcomes can be compared against a sample evaluated by people.
It is used to quickly test new model versions, to compare outputs of different systems at scale, and to build continuous evaluation pipelines during AI product development. It should still be used carefully, since the judge can inherit the same limitations or biases as the models it evaluates.
It is a practice that emerged as large language models matured, when their text comprehension capabilities became reliable enough to be employed in evaluation tasks as well.
LLM-as-judge è una tecnica di valutazione in cui un modello linguistico viene usato per giudicare la qualità delle risposte generate da un altro modello, al posto o in aggiunta a valutatori umani. Il modello giudice riceve la domanda originale, la risposta da valutare e criteri espliciti, e restituisce un punteggio o un confronto tra alternative. È diventata popolare perché permette di valutare grandi volumi di output senza il costo e i tempi della revisione umana.
In pratica si fornisce al modello giudice un prompt strutturato con le istruzioni di valutazione, spesso accompagnato da esempi di riferimento. Il modello analizza la risposta secondo criteri come correttezza, chiarezza o aderenza alle istruzioni e produce un giudizio motivato. Per aumentare l'affidabilità si possono usare più modelli giudice o confrontare i loro esiti con un campione valutato da persone.
È utilizzata per testare rapidamente nuove versioni di un modello, per confrontare output di sistemi diversi su larga scala e per creare pipeline di valutazione continua durante lo sviluppo di un prodotto AI. Va comunque usata con attenzione, perché il giudice può ereditare gli stessi limiti o pregiudizi dei modelli che valuta.
È una pratica emersa con la maturazione dei modelli linguistici di grandi dimensioni, quando le loro capacità di comprensione del testo sono diventate sufficientemente affidabili da essere impiegate anche in compiti di valutazione.
From our network
Kaimaki Web — Websites That Win Customers
Custom websites, web apps and digital marketing for growing businesses.
Visit kaimakiweb.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.