AI Dictionary › Fondamenti AI
Warm-up (del learning rate)
Warm-up is a training technique in which the learning rate starts at a very low value and is gradually increased in the early stages of training, before following its normal schedule. It prevents the model from receiving overly abrupt updates while its weights are still random and unstable. It is a common practice when training deep neural networks and large language models.
Warm-up is a training technique in which the learning rate starts at a very low value and is gradually increased in the early stages of training, before following its normal schedule. It prevents the model from receiving overly abrupt updates while its weights are still random and unstable. It is a common practice when training deep neural networks and large language models.
At the start of training, computed gradients can be very noisy, since the model has not yet learned anything. A high learning rate at this stage risks causing unstable updates or divergence. Warm-up increases the learning rate linearly or gradually over a set number of initial steps, giving the model time to stabilize before proceeding at full speed.
It is widely used when training transformers and other deep architectures, often combined with a later learning-rate decay phase. It is one of the hyperparameters configured in the training schedule, along with the number of warm-up steps and the maximum learning rate value.
It is an established engineering practice in deep neural network training, becoming particularly relevant as model sizes grew and adaptive optimizers became widely adopted.
Il warm-up è una tecnica di addestramento in cui il tasso di apprendimento parte da un valore molto basso e viene aumentato gradualmente nelle prime fasi del training, prima di seguire il suo andamento normale. Serve a evitare che il modello riceva aggiornamenti troppo bruschi quando i suoi pesi sono ancora casuali e instabili. È una pratica comune nell'addestramento di reti neurali profonde e di grandi modelli linguistici.
All'inizio dell'addestramento i gradienti calcolati possono essere molto rumorosi, perché il modello non ha ancora imparato nulla. Un tasso di apprendimento alto in questa fase rischia di causare aggiornamenti instabili o divergenza. Il warm-up aumenta il tasso di apprendimento in modo lineare o graduale per un certo numero di passi iniziali, dando al modello il tempo di stabilizzarsi prima di procedere a piena velocità.
È largamente usato nell'addestramento dei transformer e di altre architetture profonde, spesso combinato con una fase successiva di decadimento del tasso di apprendimento. È uno degli iperparametri configurati nella pianificazione dell'addestramento, insieme al numero di passi di warm-up e al valore massimo del tasso di apprendimento.
È una pratica ingegneristica consolidata nell'addestramento di reti neurali profonde, diventata particolarmente rilevante con l'aumento delle dimensioni dei modelli e con l'adozione diffusa degli ottimizzatori adattivi.
From our network
HSE Genius — AI for Safety Data Sheets
Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.
Visit hsegenius.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.