AI Dictionary › Modelli AI

Cross-Entropy

Entropia incrociata

Cross-entropy is a mathematical function used to measure how far a model's predicted probability distribution deviates from the correct expected distribution, and it is the loss function most commonly used in classification problems, including training language models to predict the next token.

Definition

What it is

Cross-entropy is a mathematical function used to measure how far a model's predicted probability distribution deviates from the correct expected distribution, and it is the loss function most commonly used in classification problems, including training language models to predict the next token.

How it works

When a model produces, for example, the probability that the next token is each word in the vocabulary, cross-entropy compares this distribution with the "correct" one, in which the token actually observed has maximum probability and all others have zero probability. The resulting value is lower the more probability the model assigns to the correct token, and higher the more confidently the model gets it wrong.

Applications

It is the loss function used in training nearly all modern language models during pre-training, where at each step the model is trained to minimize the cross-entropy between its prediction and the token actually present in the training text. It also underlies the perplexity measure used to evaluate the quality of a language model.

History & etymology

The term derives from the concept of entropy in information theory, which measures the average uncertainty of a probability distribution; "cross" entropy extends this idea to comparing two different distributions, and it was adopted in machine learning as a natural measure of the distance between a model's predictions and observed reality.

Definizione (italiano)

L'entropia incrociata è una funzione matematica usata per misurare quanto la distribuzione di probabilità prevista da un modello si discosti dalla distribuzione corretta attesa, ed è la funzione di perdita più comunemente usata nei problemi di classificazione, incluso l'addestramento dei modelli linguistici a prevedere il token successivo.

Quando un modello produce, per esempio, la probabilità che il prossimo token sia ciascuna delle parole del vocabolario, l'entropia incrociata confronta questa distribuzione con quella "corretta", in cui il token effettivamente osservato ha probabilità massima e tutti gli altri probabilità nulla. Il valore risultante è tanto più basso quanto più il modello assegna alta probabilità al token corretto, e tanto più alto quanto più il modello sbaglia con sicurezza.

È la funzione di perdita usata nell'addestramento della quasi totalità dei modelli linguistici moderni durante la fase di pre-addestramento, dove a ogni passo il modello viene addestrato a minimizzare l'entropia incrociata tra la sua previsione e il token effettivamente presente nel testo di addestramento. È inoltre alla base della misura di perplessità usata per valutare la qualità di un modello linguistico.

Il termine deriva dal concetto di entropia della teoria dell'informazione, che misura l'incertezza media di una distribuzione di probabilità; l'entropia "incrociata" estende questa idea al confronto tra due distribuzioni diverse, ed è stata adottata nel machine learning come misura naturale della distanza tra previsioni del modello e realtà osservata.

Related terms

More in Modelli AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.