AI Dictionary › Fondamenti AI
Denial of Service su LLM
LLM denial of service is a category of attack in which an attacker seeks to degrade or interrupt the availability of a service based on a language model, not necessarily by stealing data or manipulating output, but simply by exhausting its computational or economic resources. Unlike classic denial of service against web servers, which relies on saturating bandwidth or network connections, attacks against LLMs exploit the specific characteristics of inference computation.
LLM denial of service is a category of attack in which an attacker seeks to degrade or interrupt the availability of a service based on a language model, not necessarily by stealing data or manipulating output, but simply by exhausting its computational or economic resources. Unlike classic denial of service against web servers, which relies on saturating bandwidth or network connections, attacks against LLMs exploit the specific characteristics of inference computation.
Techniques include sending prompts crafted to maximize processing time or the length of the generated output, potentially exploiting longer reasoning cycles in models with extended reasoning capabilities; sending a high volume of legitimate but useless requests to saturate available capacity and drive up compute costs or response times for other users; or prompts that, exploiting specific model weaknesses, induce it into repetitive loops that consume tokens without producing useful output.
It is an economic risk as well as a technical one for anyone offering a pay-as-you-go AI service: since language model APIs are typically billed based on tokens processed, an attack of this kind can translate into significant direct costs on top of service degradation for legitimate users. Main defenses include rate limits, maximum limits on input and output length, and monitoring of anomalous usage patterns.
The concept adapts to the language model domain a well-known cyberattack category dating back to the 1990s in the context of networks and web servers; the specific application to LLMs emerged with the spread of paid commercial APIs starting in 2022-2023.
Il denial of service su LLM è una categoria di attacco in cui un aggressore cerca di degradare o interrompere la disponibilità di un servizio basato su un modello linguistico, non necessariamente rubando dati o manipolando l'output, ma semplicemente esaurendone le risorse computazionali o economiche. A differenza del denial of service classico contro server web, che si basa sul saturare la banda o le connessioni di rete, quello contro gli LLM sfrutta le caratteristiche specifiche del calcolo di inferenza.
Le tecniche includono l'invio di prompt costruiti per massimizzare il tempo di elaborazione o la lunghezza dell'output generato, sfruttando magari cicli di ragionamento più lunghi nei modelli con capacità di reasoning esteso; l'invio di un volume elevato di richieste legittime ma inutili per saturare la capacità disponibile e far salire i costi di calcolo o i tempi di risposta per gli altri utenti; oppure prompt che, sfruttando debolezze specifiche del modello, lo inducono a entrare in cicli ripetitivi che consumano token senza produrre un output utile.
È un rischio economico oltre che tecnico per chi offre un servizio AI a consumo: dato che le API dei modelli linguistici sono tipicamente fatturate in base ai token elaborati, un attacco di questo tipo può tradursi in costi diretti significativi oltre che nel degrado del servizio per gli utenti legittimi. Le difese principali includono limiti di frequenza, limiti massimi sulla lunghezza di input e output, e monitoraggio dei pattern di utilizzo anomali.
Il concetto adatta al dominio dei modelli linguistici una categoria di attacco informatico ben nota fin dagli anni '90 nel contesto delle reti e dei server web; l'applicazione specifica agli LLM è emersa con la diffusione delle API commerciali a pagamento a partire dal 2022-2023.
From our network
AGORÀ Intelligence — Enterprise AI Governance Platform
Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.
Visit agora-intelligence.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.