AI Dictionary › Fondamenti AI

Leaderboard

Classifica di valutazione (Leaderboard)

A leaderboard is a public ranking that orders AI models by the scores they achieve on one or more benchmarks. It lets researchers, companies and developers quickly compare the performance of different models on common tasks, such as language understanding, mathematical reasoning or code generation. It is a widely used reference point for navigating the dozens of available models.

Definition

What it is

A leaderboard is a public ranking that orders AI models by the scores they achieve on one or more benchmarks. It lets researchers, companies and developers quickly compare the performance of different models on common tasks, such as language understanding, mathematical reasoning or code generation. It is a widely used reference point for navigating the dozens of available models.

How it works

Each leaderboard defines its own rules: which tests it includes, how scores are weighted, and how often it is updated. Some rely on standardized automated tests, others incorporate preference votes collected from real users comparing responses from different models. Scores are recalculated as new models or new versions of existing ones are released.

Applications

Leaderboards are consulted by those choosing a model for a product, by researchers positioning their work against the state of the art, and by teams communicating a new model's progress to the public. They have become a common reference point in debates over which system is most advanced at a given moment.

History & etymology

AI leaderboards spread alongside the proliferation of language models, when it became useful to have shared, updatable reference points instead of isolated comparisons between just two systems.

Definizione (italiano)

Una leaderboard è una classifica pubblica che ordina i modelli AI in base ai punteggi ottenuti su uno o più benchmark. Permette a ricercatori, aziende e sviluppatori di confrontare rapidamente le prestazioni di modelli diversi su compiti comuni, come comprensione del linguaggio, ragionamento matematico o generazione di codice. È uno strumento di riferimento molto usato per orientarsi tra le decine di modelli disponibili.

Ogni leaderboard definisce le proprie regole: quali test include, come vengono pesati i punteggi e con quale frequenza viene aggiornata. Alcune si basano su test automatici standardizzati, altre integrano voti di preferenza raccolti da utenti reali che confrontano risposte di modelli diversi. I punteggi vengono ricalcolati man mano che escono nuovi modelli o nuove versioni di quelli esistenti.

Le leaderboard vengono consultate da chi deve scegliere un modello per un prodotto, da ricercatori che vogliono posizionare il proprio lavoro rispetto allo stato dell'arte e da chi comunica i progressi di un nuovo modello al pubblico. Sono diventate un riferimento comune nel dibattito su quale sistema sia più avanzato in un dato momento.

Le leaderboard per l'AI si sono diffuse insieme alla proliferazione dei modelli linguistici, quando è diventato utile avere punti di riferimento condivisi e aggiornabili invece di confronti isolati tra due soli sistemi.

Related terms

More in Fondamenti AI

Put it into practice

From our network

Kaimaki Web — Websites That Win Customers

Custom websites, web apps and digital marketing for growing businesses.

Visit kaimakiweb.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.