AI Dictionary › AI Models
LLM come giudice (LLM-as-Judge)
LLM-as-judge is an evaluation technique in which a language model is used to judge the quality of responses generated by another model, in place of or alongside human evaluators. The judge model receives the original question, the response to be evaluated and explicit criteria, and returns a score or a comparison between alternatives. It has become popular because it allows large volumes of output to be evaluated without the cost and time of human review.
In practice, the judge model is given a structured prompt with evaluation instructions, often accompanied by reference examples. The model analyzes the response according to criteria such as correctness, clarity or adherence to instructions, and produces a reasoned judgment. To increase reliability, multiple judge models can be used, or their outcomes can be compared against a sample evaluated by people.
It is used to quickly test new model versions, to compare outputs of different systems at scale, and to build continuous evaluation pipelines during AI product development. It should still be used carefully, since the judge can inherit the same limitations or biases as the models it evaluates.
It is a practice that emerged as large language models matured, when their text comprehension capabilities became reliable enough to be employed in evaluation tasks as well.
Grace is experimenting with LLM-as-judge for a first automatic evaluation of open-ended answers in scenarios, ahead of possible human review of the most uncertain cases.
From our network
Kaimaki Web: Websites That Win Customers
Custom websites, web apps and digital marketing for growing businesses.
Visit kaimakiweb.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.