AI Dictionary › AI Fundamentals
Weight decay is a regularization technique that penalizes overly large weights during training, adding a term proportional to their magnitude to the loss function. The goal is to keep the model's weights contained, favoring simpler solutions and reducing the risk of overfitting. It is one of the most common hyperparameters configured when training neural networks.
During each update, in addition to pushing weights toward values that reduce error on the data, the optimizer slightly shrinks their magnitude, like a constant force pulling them toward zero. The strength of this penalty is controlled by a coefficient that balances fitting the data against keeping the model simple.
It is widely used when training deep neural networks, often combined with other regularization methods such as dropout and early stopping. It is built into many modern optimizers, including variants specifically designed to apply it more effectively alongside adaptive learning-rate updates.
It is a regularization technique with roots in classical statistical methods for penalizing parameters, long adopted in neural network training.
From our network
INDACO TMS: Transport Management for European Logistics
Shipment tracking, multi-carrier EDI and automated invoicing in one cloud platform. Invoices generated in under 10 seconds.
Visit indacotms.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.