AI Dictionary › AI Models
Entropia incrociata
Cross-entropy is a mathematical function used to measure how far a model's predicted probability distribution deviates from the correct expected distribution, and it is the loss function most commonly used in classification problems, including training language models to predict the next token.
When a model produces, for example, the probability that the next token is each word in the vocabulary, cross-entropy compares this distribution with the "correct" one, in which the token actually observed has maximum probability and all others have zero probability. The resulting value is lower the more probability the model assigns to the correct token, and higher the more confidently the model gets it wrong.
It is the loss function used in training nearly all modern language models during pre-training, where at each step the model is trained to minimize the cross-entropy between its prediction and the token actually present in the training text. It also underlies the perplexity measure used to evaluate the quality of a language model.
The term derives from the concept of entropy in information theory, which measures the average uncertainty of a probability distribution; "cross" entropy extends this idea to comparing two different distributions, and it was adopted in machine learning as a natural measure of the distance between a model's predictions and observed reality.
From our network
AGORÀ Intelligence: Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.