AI Dictionary › Fondamenti AI
Confronto a coppie (Pairwise Comparison)
Pairwise comparison is an evaluation method in which a judge, human or automated, is shown two responses generated by different models (or by the same model with different settings) for the same question, and asked which of the two is better. It is an alternative to assigning an absolute score, often simpler and more reliable because comparing two options is cognitively easier than rating one on its own on a numeric scale.
In practice, many pairs of responses across different questions are collected, and the preference expressed for each comparison is recorded. Results are then aggregated using statistical models, similar to those used in sports ranking systems, to produce an overall ranking of the models involved. The more comparisons collected, the more stable the resulting ranking becomes.
It is widely used in public language model leaderboards, where users vote on which of two anonymous responses they prefer, and in model alignment processes, where the collected preferences are used to train the system to produce better responses. It is also a practical tool for A/B testing between different versions of the same AI product.
It is a technique drawn from psychometrics and choice theory, where pairwise comparisons have long been used to build preference scales from relative judgments, which are simpler to express than absolute scores.
Grace uses pairwise comparisons in internal testing, showing evaluators two alternative answers to the same scenario to calibrate the scoring rubrics.
From our network
AGORÀ Intelligence: Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.