AI Dictionary › Fondamenti AI

Pairwise Comparison

Confronto a coppie (Pairwise Comparison)

Pairwise comparison is an evaluation method in which a judge, human or automated, is shown two responses generated by different models (or by the same model with different settings) for the same question, and asked which of the two is better. It is an alternative to assigning an absolute score, often simpler and more reliable because comparing two options is cognitively easier than rating one on its own on a numeric scale.

Definition

What it is

Pairwise comparison is an evaluation method in which a judge, human or automated, is shown two responses generated by different models (or by the same model with different settings) for the same question, and asked which of the two is better. It is an alternative to assigning an absolute score, often simpler and more reliable because comparing two options is cognitively easier than rating one on its own on a numeric scale.

How it works

In practice, many pairs of responses across different questions are collected, and the preference expressed for each comparison is recorded. Results are then aggregated using statistical models, similar to those used in sports ranking systems, to produce an overall ranking of the models involved. The more comparisons collected, the more stable the resulting ranking becomes.

Applications

It is widely used in public language model leaderboards, where users vote on which of two anonymous responses they prefer, and in model alignment processes, where the collected preferences are used to train the system to produce better responses. It is also a practical tool for A/B testing between different versions of the same AI product.

History & etymology

It is a technique drawn from psychometrics and choice theory, where pairwise comparisons have long been used to build preference scales from relative judgments, which are simpler to express than absolute scores.

Definizione (italiano)

Il confronto a coppie è un metodo di valutazione in cui a un giudice, umano o automatico, vengono mostrate due risposte generate da modelli diversi (o dallo stesso modello con impostazioni diverse) per la stessa domanda, chiedendo quale delle due sia migliore. È un'alternativa all'assegnazione di un punteggio assoluto, spesso più semplice e affidabile perché confrontare due opzioni è cognitivamente più facile che valutarne una da sola su una scala numerica.

In pratica si raccolgono molte coppie di risposte su domande diverse e si registra la preferenza espressa per ciascun confronto. I risultati vengono poi aggregati con modelli statistici, simili a quelli usati nei sistemi di ranking sportivi, per ottenere una classifica complessiva dei modelli coinvolti. Più confronti vengono raccolti, più stabile diventa la classifica risultante.

È molto usato nelle leaderboard pubbliche di modelli linguistici, dove gli utenti votano quale tra due risposte anonime preferiscono, e nei processi di allineamento dei modelli, dove le preferenze raccolte servono ad addestrare il sistema a produrre risposte migliori. È anche uno strumento pratico per test A/B tra versioni diverse di uno stesso prodotto AI.

È una tecnica che deriva dalla psicometria e dalla teoria delle scelte, dove i confronti a coppie sono usati da molto tempo per costruire scale di preferenza a partire da giudizi relativi più semplici da esprimere rispetto a punteggi assoluti.

Related terms

More in Fondamenti AI

Put it into practice

From our network

INDACO TMS — Transport Management for European Logistics

Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.

Visit indacotms.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.