AI Dictionary › Fondamenti AI
Fuori distribuzione (Out-of-Distribution)
Out-of-distribution refers to an input that differs significantly from the data an AI model was trained on, in content, format or context. On such inputs, a model's performance is generally less reliable, since the system faces patterns it never encountered during training. Recognizing these cases matters just as much as measuring accuracy on familiar data.
Out-of-distribution refers to an input that differs significantly from the data an AI model was trained on, in content, format or context. On such inputs, a model's performance is generally less reliable, since the system faces patterns it never encountered during training. Recognizing these cases matters just as much as measuring accuracy on familiar data.
To evaluate out-of-distribution behavior, dedicated test sets are built with inputs deliberately different from typical ones, such as new domains, less-represented languages or unusual formats. Changes in the model's accuracy and reliability relative to in-distribution data are then measured, often also observing whether the model signals its own uncertainty instead of answering with excessive confidence. Good out-of-distribution behavior includes the ability to recognize its own limits.
It is a central concern for applications where real-world inputs can vary widely from training data, such as analyzing domain-specific documents or use in new markets and languages. It also helps decide when a system should decline to answer or request human intervention.
The concept originates in statistics and classical machine learning, where the distinction between in-distribution and out-of-distribution data has long been used to assess the generalization ability of predictive models.
Out-of-distribution indica un input che si discosta significativamente dai dati su cui un modello AI è stato addestrato, per contenuto, formato o contesto. Su questi input le prestazioni di un modello sono generalmente meno affidabili, perché il sistema si trova ad affrontare pattern mai visti durante l'addestramento. Riconoscere questi casi è importante tanto quanto misurare l'accuratezza sui dati abituali.
Per valutare il comportamento fuori distribuzione si creano appositi set di test con input volutamente diversi da quelli tipici, come domini nuovi, lingue meno rappresentate o formati insoliti. Si misura poi come cambiano accuratezza e affidabilità del modello rispetto ai dati in distribuzione, spesso osservando anche se il modello segnala la propria incertezza invece di rispondere con eccessiva sicurezza. Un buon comportamento fuori distribuzione include la capacità di riconoscere i propri limiti.
È un tema centrale per applicazioni dove gli input reali possono variare molto rispetto ai dati di addestramento, come l'analisi di documenti settoriali specifici o l'uso in nuovi mercati e lingue. Serve anche a decidere quando un sistema dovrebbe rifiutarsi di rispondere o richiedere l'intervento umano.
Il concetto nasce nella statistica e nell'apprendimento automatico classico, dove la distinzione tra dati in distribuzione e fuori distribuzione è usata da tempo per valutare la capacità di generalizzazione dei modelli predittivi.
From our network
HSE Genius — AI for Safety Data Sheets
Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.
Visit hsegenius.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.