AI Dictionary › Fondamenti AI

Training Data Extraction

Estrazione dei Dati di Addestramento

Training data extraction is a technique by which an attacker induces a language model to reproduce verbatim fragments of text memorized during training, rather than generating new, generalized text. It is the most direct risk tied to unintentional memorization: if a model has seen a certain document enough times, or if that document was particularly distinctive, it may repeat it almost identically when given a prompt that recalls its beginning.

Definition

What it is

Training data extraction is a technique by which an attacker induces a language model to reproduce verbatim fragments of text memorized during training, rather than generating new, generalized text. It is the most direct risk tied to unintentional memorization: if a model has seen a certain document enough times, or if that document was particularly distinctive, it may repeat it almost identically when given a prompt that recalls its beginning.

How it works

Technically the attack often works by feeding the model an initial fragment of a text suspected to be in the training set and observing whether the model completes the rest too precisely to be coincidental, or by repeating the same request many times with small variations to increase the chance of surfacing a memorized fragment. Larger models, paradoxically, tend to memorize rare or distinctive text passages more easily than smaller ones.

Applications

It is a concrete risk when training data includes personal information, proprietary code, or copyrighted material: a model could return, verbatim, data that should never appear in output, such as phone numbers, email addresses, or entire passages from copyrighted books. Main defenses include deduplicating training data, which reduces the repetitions that favor memorization, and differential privacy techniques that mathematically limit how much a single document can imprint itself on the model.

History & etymology

The phenomenon was publicly and systematically documented by the research community starting in 2020-2021, when experiments on large language models showed it was possible to extract verbatim memorized text with relatively simple requests, prompting AI labs to invest in specific mitigations.

Definizione (italiano)

L'estrazione dei dati di addestramento è una tecnica con cui un aggressore induce un modello linguistico a riprodurre alla lettera frammenti di testo memorizzati durante l'addestramento, anziché generare testo nuovo e generalizzato. È il rischio più diretto legato alla memorizzazione involontaria: se un modello ha visto un certo documento abbastanza volte, o se quel documento era particolarmente distintivo, può capitare che lo ripeta quasi identico quando riceve un prompt che ne ricorda l'inizio.

Tecnicamente l'attacco funziona spesso fornendo al modello un frammento iniziale di un testo sospettato di essere nel training set e osservando se il modello ne completa il resto in modo troppo preciso per essere casuale, oppure ripetendo la stessa richiesta molte volte con piccole variazioni per aumentare la probabilità di far emergere un frammento memorizzato. I modelli più grandi, paradossalmente, tendono a memorizzare più facilmente porzioni di testo rare o distintive rispetto a modelli più piccoli.

È un rischio concreto quando i dati di addestramento includono informazioni personali, codice proprietario o materiale protetto da copyright: un modello potrebbe restituire, testualmente, dati che non dovrebbero mai comparire nell'output, come numeri di telefono, indirizzi email o passaggi interi di libri sotto copyright. Le difese principali includono la deduplicazione dei dati di addestramento, che riduce le ripetizioni che favoriscono la memorizzazione, e tecniche di privacy differenziale che limitano matematicamente quanto un singolo documento possa imprimersi nel modello.

Il fenomeno è stato documentato pubblicamente e in modo sistematico dalla comunità di ricerca a partire dal 2020-2021, quando esperimenti su modelli linguistici di grandi dimensioni hanno dimostrato che era possibile estrarre testo memorizzato verbatim con richieste relativamente semplici, portando i laboratori AI a investire in mitigazioni specifiche.

Related terms

More in Fondamenti AI

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.