AI Dictionary › Fondamenti AI
Classifica di valutazione (Leaderboard)
A leaderboard is a public ranking that orders AI models by the scores they achieve on one or more benchmarks. It lets researchers, companies and developers quickly compare the performance of different models on common tasks, such as language understanding, mathematical reasoning or code generation. It is a widely used reference point for navigating the dozens of available models.
A leaderboard is a public ranking that orders AI models by the scores they achieve on one or more benchmarks. It lets researchers, companies and developers quickly compare the performance of different models on common tasks, such as language understanding, mathematical reasoning or code generation. It is a widely used reference point for navigating the dozens of available models.
Each leaderboard defines its own rules: which tests it includes, how scores are weighted, and how often it is updated. Some rely on standardized automated tests, others incorporate preference votes collected from real users comparing responses from different models. Scores are recalculated as new models or new versions of existing ones are released.
Leaderboards are consulted by those choosing a model for a product, by researchers positioning their work against the state of the art, and by teams communicating a new model's progress to the public. They have become a common reference point in debates over which system is most advanced at a given moment.
AI leaderboards spread alongside the proliferation of language models, when it became useful to have shared, updatable reference points instead of isolated comparisons between just two systems.
Una leaderboard è una classifica pubblica che ordina i modelli AI in base ai punteggi ottenuti su uno o più benchmark. Permette a ricercatori, aziende e sviluppatori di confrontare rapidamente le prestazioni di modelli diversi su compiti comuni, come comprensione del linguaggio, ragionamento matematico o generazione di codice. È uno strumento di riferimento molto usato per orientarsi tra le decine di modelli disponibili.
Ogni leaderboard definisce le proprie regole: quali test include, come vengono pesati i punteggi e con quale frequenza viene aggiornata. Alcune si basano su test automatici standardizzati, altre integrano voti di preferenza raccolti da utenti reali che confrontano risposte di modelli diversi. I punteggi vengono ricalcolati man mano che escono nuovi modelli o nuove versioni di quelli esistenti.
Le leaderboard vengono consultate da chi deve scegliere un modello per un prodotto, da ricercatori che vogliono posizionare il proprio lavoro rispetto allo stato dell'arte e da chi comunica i progressi di un nuovo modello al pubblico. Sono diventate un riferimento comune nel dibattito su quale sistema sia più avanzato in un dato momento.
Le leaderboard per l'AI si sono diffuse insieme alla proliferazione dei modelli linguistici, quando è diventato utile avere punti di riferimento condivisi e aggiornabili invece di confronti isolati tra due soli sistemi.
From our network
Kaimaki Web — Websites That Win Customers
Custom websites, web apps and digital marketing for growing businesses.
Visit kaimakiweb.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.