AI Dictionary › AI Fundamentals
Attacco di Estrazione del Modello
A model extraction attack is a technique by which an attacker tries to reconstruct a proprietary AI model, or a sufficiently accurate approximation of it, by repeatedly querying it through its public interface and observing the outputs produced. The goal is not to gain direct access to the model's weights, which remain protected, but to train a substitute model that replicates its behavior without paying its development costs.
Technically, the attacker sends a large volume of systematically varied requests, collects the resulting input-output pairs, and uses them as a training dataset for a new model, in a process similar to distillation but conducted without the original model owner's consent. The more queries the attacker manages to send and the more informative the responses obtained, the more faithful the extracted model turns out.
It is a concrete risk for companies offering proprietary AI models through paid APIs, where the commercial value lies precisely in the model's uniqueness: a competitor able to extract a cheap approximation would erode the competitive advantage without having borne the original training costs. Defenses include rate limits on requests, controlled obfuscation of probabilities returned alongside the output, and monitoring of query patterns indicating a systematic extraction attempt rather than normal use.
The concept originated in machine learning security research as early as the mid-2010s, initially applied to simpler classification models offered as a cloud service, and naturally extended to large language models with the spread of commercial APIs starting in 2020.
From our network
Magellano GPS: Fleet Tracking Made Simple
Real-time GPS tracking, remote engine lock, fuel and CO₂ reporting for your fleet.
Visit magellanogps.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.