AI Dictionary › Fondamenti AI
Contaminazione dei dati di test (Test Set Contamination)
Test set contamination occurs when the content used to evaluate an AI model was already present, in whole or in part, in the data the model was trained on. In this case, the model is not demonstrating genuine generalization ability but is simply reproducing information it has already seen, artificially inflating the score obtained. It is one of the most insidious problems in evaluating large-scale language models.
Test set contamination occurs when the content used to evaluate an AI model was already present, in whole or in part, in the data the model was trained on. In this case, the model is not demonstrating genuine generalization ability but is simply reproducing information it has already seen, artificially inflating the score obtained. It is one of the most insidious problems in evaluating large-scale language models.
It can happen because training data, collected from very broad public sources such as the web, unintentionally includes the same texts used as benchmarks. To detect it, texts in the evaluation set are compared with those in the training set, looking for direct or partial overlaps. When contamination is found, scores obtained on that benchmark are considered unreliable.
It is a central concern when new models are released and their results on public benchmarks are communicated, because a score inflated by contamination can lead to wrong choices about which model to adopt. For this reason, benchmarks kept private or refreshed periodically are becoming more common, precisely to reduce the risk of them ending up in future training data.
It is a problem that emerged with the enormous growth of datasets used to train language models, when it became difficult to guarantee that evaluation texts remained completely separate from training texts.
La contaminazione dei dati di test si verifica quando i contenuti usati per valutare un modello AI erano già presenti, in tutto o in parte, nei dati con cui il modello è stato addestrato. In questo caso il modello non sta dimostrando una vera capacità di generalizzazione, ma sta semplicemente riproducendo informazioni già viste, gonfiando artificialmente il punteggio ottenuto. È uno dei problemi più insidiosi nella valutazione dei modelli linguistici di grande scala.
Può accadere perché i dati di addestramento, raccolti da fonti pubbliche molto ampie come il web, includono involontariamente gli stessi testi usati come benchmark. Per individuarla si confrontano i testi del set di valutazione con quelli del set di addestramento cercando sovrapposizioni dirette o parziali. Quando la contaminazione viene rilevata, i punteggi ottenuti su quel benchmark vengono considerati inaffidabili.
È un tema centrale quando si pubblicano nuovi modelli e si comunicano i loro risultati sui benchmark pubblici, perché un punteggio gonfiato dalla contaminazione può indurre a scelte sbagliate su quale modello adottare. Per questo motivo si stanno diffondendo benchmark tenuti privati o rinnovati periodicamente, proprio per ridurre il rischio che finiscano nei dati di addestramento futuri.
È un problema emerso con la crescita smisurata dei dataset usati per addestrare i modelli linguistici, quando è diventato difficile garantire che i testi di valutazione restassero completamente separati da quelli di addestramento.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.