AI Dictionary › Fondamenti AI

Model Extraction Attack

Attacco di Estrazione del Modello

A model extraction attack is a technique by which an attacker tries to reconstruct a proprietary AI model, or a sufficiently accurate approximation of it, by repeatedly querying it through its public interface and observing the outputs produced. The goal is not to gain direct access to the model's weights, which remain protected, but to train a substitute model that replicates its behavior without paying its development costs.

Definition

What it is

A model extraction attack is a technique by which an attacker tries to reconstruct a proprietary AI model, or a sufficiently accurate approximation of it, by repeatedly querying it through its public interface and observing the outputs produced. The goal is not to gain direct access to the model's weights, which remain protected, but to train a substitute model that replicates its behavior without paying its development costs.

How it works

Technically, the attacker sends a large volume of systematically varied requests, collects the resulting input-output pairs, and uses them as a training dataset for a new model, in a process similar to distillation but conducted without the original model owner's consent. The more queries the attacker manages to send and the more informative the responses obtained, the more faithful the extracted model turns out.

Applications

It is a concrete risk for companies offering proprietary AI models through paid APIs, where the commercial value lies precisely in the model's uniqueness: a competitor able to extract a cheap approximation would erode the competitive advantage without having borne the original training costs. Defenses include rate limits on requests, controlled obfuscation of probabilities returned alongside the output, and monitoring of query patterns indicating a systematic extraction attempt rather than normal use.

History & etymology

The concept originated in machine learning security research as early as the mid-2010s, initially applied to simpler classification models offered as a cloud service, and naturally extended to large language models with the spread of commercial APIs starting in 2020.

Definizione (italiano)

Un attacco di estrazione del modello è una tecnica con cui un aggressore cerca di ricostruire un modello AI proprietario, o un'approssimazione sufficientemente accurata di esso, interrogandolo ripetutamente tramite la sua interfaccia pubblica e osservando gli output prodotti. Lo scopo non è ottenere l'accesso diretto ai pesi del modello, che restano protetti, ma addestrare un modello sostituto che ne replica il comportamento senza pagarne i costi di sviluppo.

Tecnicamente l'attaccante invia un grande volume di richieste sistematicamente variate, raccoglie le coppie input-output risultanti e le usa come dataset di addestramento per un nuovo modello, in un processo simile alla distillazione ma condotto senza il consenso del proprietario del modello originale. Più query l'aggressore riesce a inviare e più informative sono le risposte ottenute, più fedele risulta il modello estratto.

È un rischio concreto per le aziende che offrono modelli AI proprietari tramite API a pagamento, dove il valore commerciale risiede proprio nell'unicità del modello: un concorrente che riuscisse a estrarne un'approssimazione economica eroderebbe il vantaggio competitivo senza aver sostenuto i costi di addestramento originali. Le difese includono limiti di frequenza sulle richieste, offuscamento controllato delle probabilità restituite insieme all'output, e monitoraggio dei pattern di query che indicano un tentativo sistematico di estrazione anziché un uso normale.

Il concetto nasce nella ricerca sulla sicurezza del machine learning già a metà degli anni 2010, applicato inizialmente a modelli di classificazione più semplici offerti come servizio cloud, e si è esteso naturalmente ai grandi modelli linguistici con la diffusione delle API commerciali a partire dal 2020.

Related terms

More in Fondamenti AI

Put it into practice

From our network

Magellano GPS — Fleet Tracking Made Simple

Real-time GPS tracking, remote engine lock, fuel and CO₂ reporting for your fleet.

Visit magellanogps.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.