AI Dictionary › Fondamenti AI
A checkpoint is a saved snapshot of a model's state at a given point during training, including the weights learned up to that moment. It works like a save point that allows training to be resumed, the model to be shared, or it to be used directly for inference. Without checkpoints, any interruption to training would mean losing all progress made.
A checkpoint is a saved snapshot of a model's state at a given point during training, including the weights learned up to that moment. It works like a save point that allows training to be resumed, the model to be shared, or it to be used directly for inference. Without checkpoints, any interruption to training would mean losing all progress made.
During training, the process periodically saves the state of the parameters, often at regular time intervals or after a certain number of epochs. Each checkpoint can be loaded to resume training from that point, to compare performance across different phases, or to select the version with the best results on validation data.
Checkpoints make long, expensive training runs more resilient, protecting against hardware failures or unexpected interruptions. They are also the format in which pre-trained models are publicly distributed, letting others start from an already-trained state instead of from scratch.
The term comes from the general computing practice of saving an intermediate program state, later applied specifically to machine learning model training.
Un checkpoint è un salvataggio dello stato di un modello a un certo punto dell'addestramento, comprensivo dei pesi appresi fino a quel momento. Funziona come un'istantanea che permette di riprendere il lavoro, condividere il modello o usarlo direttamente per l'inferenza. Senza checkpoint, ogni interruzione dell'addestramento comporterebbe la perdita di tutto il progresso fatto.
Durante l'addestramento, il processo salva periodicamente lo stato dei parametri, spesso a intervalli regolari di tempo o dopo un certo numero di epoche. Ogni checkpoint può essere caricato per riprendere l'addestramento da quel punto, per confrontare le prestazioni tra fasi diverse o per selezionare la versione con i risultati migliori sui dati di validazione.
I checkpoint sono usati per rendere robusti training molto lunghi e costosi, proteggendo dal rischio di guasti hardware o interruzioni impreviste. Sono anche il formato con cui i modelli pre-addestrati vengono distribuiti pubblicamente, permettendo ad altri di partire da uno stato già addestrato invece che da zero.
Il termine deriva dalla pratica generale dell'informatica di salvare uno stato intermedio di un programma, applicata poi in modo specifico all'addestramento dei modelli di machine learning.
From our network
Kaimaki Web — Websites That Win Customers
Custom websites, web apps and digital marketing for growing businesses.
Visit kaimakiweb.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.