AI Dictionary › Prompting

Adversarial Prompt

Prompt Avversariale

An adversarial prompt is an input deliberately crafted to confuse, deceive, or force an AI model to behave differently from what its safety constraints intend. The term is the umbrella under which many specific techniques fall — jailbreak, prompt injection, prompt leaking — united by the intent to exploit the model's weaknesses rather than use it as designed.

Definition

What it is

An adversarial prompt is an input deliberately crafted to confuse, deceive, or force an AI model to behave differently from what its safety constraints intend. The term is the umbrella under which many specific techniques fall — jailbreak, prompt injection, prompt leaking — united by the intent to exploit the model's weaknesses rather than use it as designed.

How it works

Adversarial attacks on language models exploit the statistical nature of responses: ambiguous phrasing, role-play framing, encapsulating malicious text inside innocuous contexts (translations, poems, code), or character sequences designed to confuse the tokenizer. Unlike classic adversarial attacks on image neural networks, which alter pixels imperceptible to the human eye, those on LLMs work almost entirely at the natural language level, making them harder to catch with simple automated filters.

Applications

In modern AI, resistance to adversarial prompts is a standard evaluation criterion before a model ships: red-teaming teams systematically generate thousands of variants to measure how easily protections can be bypassed. Public safety benchmarks also include suites of adversarial prompts to compare different models.

History & etymology

The concept of the adversarial example originated in machine learning applied to computer vision, before large language models became widespread; with the arrival of ChatGPT and similar systems the term naturally extended into the domain of text prompting.

Definizione (italiano)

Un prompt avversariale è un input costruito deliberatamente per confondere, ingannare o forzare un modello AI a comportarsi in modo diverso da quello previsto dai suoi vincoli di sicurezza. Il termine è l'ombrello sotto cui rientrano molte tecniche specifiche — jailbreak, prompt injection, prompt leaking — accomunate dall'intento di sfruttare i punti deboli del modello anziché usarlo nel modo previsto.

Gli attacchi avversariali sui modelli linguistici sfruttano la natura statistica delle risposte: frasi ambigue, giochi di ruolo, incapsulamento del testo malevolo dentro contesti innocui (traduzioni, poesie, codice), o sequenze di caratteri studiate per confondere il tokenizzatore. A differenza degli attacchi avversariali classici sulle reti neurali per immagini, che modificano pixel impercettibili all'occhio umano, quelli sui LLM lavorano quasi sempre a livello di linguaggio naturale, il che li rende più difficili da rilevare con semplici filtri automatici.

Nell'AI moderna la resistenza ai prompt avversariali è un criterio di valutazione standard prima del rilascio di un modello: i team di red teaming generano sistematicamente migliaia di varianti per misurare quanto facilmente le protezioni vengano aggirate. Anche i benchmark di sicurezza pubblici includono suite di prompt avversariali per confrontare modelli diversi.

Il concetto di esempio avversariale nasce nel machine learning applicato alla visione artificiale, prima della diffusione dei grandi modelli linguistici; con l'arrivo di ChatGPT e sistemi simili il termine si è esteso naturalmente al dominio del prompting testuale.

Related terms

More in Prompting

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.