AI Dictionary › Fondamenti AI

Weight Decay

Weight decay is a regularization technique that penalizes overly large weights during training, adding a term proportional to their magnitude to the loss function. The goal is to keep the model's weights contained, favoring simpler solutions and reducing the risk of overfitting. It is one of the most common hyperparameters configured when training neural networks.

Definition

What it is

Weight decay is a regularization technique that penalizes overly large weights during training, adding a term proportional to their magnitude to the loss function. The goal is to keep the model's weights contained, favoring simpler solutions and reducing the risk of overfitting. It is one of the most common hyperparameters configured when training neural networks.

How it works

During each update, in addition to pushing weights toward values that reduce error on the data, the optimizer slightly shrinks their magnitude, like a constant force pulling them toward zero. The strength of this penalty is controlled by a coefficient that balances fitting the data against keeping the model simple.

Applications

It is widely used when training deep neural networks, often combined with other regularization methods such as dropout and early stopping. It is built into many modern optimizers, including variants specifically designed to apply it more effectively alongside adaptive learning-rate updates.

History & etymology

It is a regularization technique with roots in classical statistical methods for penalizing parameters, long adopted in neural network training.

Definizione (italiano)

Il weight decay è una tecnica di regolarizzazione che penalizza i pesi troppo grandi durante l'addestramento, aggiungendo alla funzione di perdita un termine proporzionale alla loro grandezza. L'obiettivo è mantenere i pesi del modello contenuti, favorendo soluzioni più semplici e riducendo il rischio di overfitting. È uno degli iperparametri più comuni da configurare nell'addestramento di reti neurali.

Durante ogni aggiornamento, oltre a spingere i pesi verso valori che riducono l'errore sui dati, l'ottimizzatore li riduce leggermente in modulo, come una forza costante che li tira verso lo zero. Il grado di questa penalizzazione è controllato da un coefficiente che bilancia l'adattamento ai dati e la semplicità del modello.

È ampiamente utilizzato nell'addestramento di reti neurali profonde, spesso in combinazione con altri metodi di regolarizzazione come il dropout e l'early stopping. È integrato in molti ottimizzatori moderni, incluse varianti specificamente progettate per applicarlo in modo più efficace insieme all'aggiornamento adattivo del tasso di apprendimento.

È una tecnica di regolarizzazione con radici nei metodi statistici classici di penalizzazione dei parametri, adottata da tempo nell'addestramento delle reti neurali.

Related terms

More in Fondamenti AI

Put it into practice

From our network

INDACO TMS — Transport Management for European Logistics

Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.

Visit indacotms.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.