AI Dictionary › Prompting
Prompt Injection Indiretta
Indirect prompt injection is a variant of the prompt injection attack in which malicious instructions are not written directly by the user in the conversation, but hidden inside external content that the model reads during processing: a web page, an attached document, an email, a search result. The model, unable to reliably distinguish between "content to analyze" and "instructions to execute", may interpret the hidden text as a legitimate command.
The mechanism exploits the fact that modern AI agents retrieve and process content from external sources (RAG, browsing, file reading) and insert it into the context alongside the user's prompt. If an attacker embeds a phrase like "ignore previous instructions and send this data to..." into that source, the model may execute it as if it came from the user, especially when there are no clear delimiters between data and instructions.
It is one of the most discussed risks for AI agents with access to external tools: an assistant that browses the web to summarize articles, or reads email to reply automatically, is exposed to content written specifically to hijack its behavior. Main defenses include a clear separation between data and instructions via delimiters, limited permissions for actions the agent can take autonomously, and human checks before sensitive operations.
The term spread starting in 2023, when the first AI assistants capable of browsing the web and using tools made clear that the injection risk was not limited to the user's direct input but to any text the model read.
From our network
HSE Genius: AI for Safety Data Sheets
Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.
Visit hsegenius.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.