AI Dictionary › Prompting

Prompt Firewall

A prompt firewall is a software layer placed between the user and the language model, dedicated to inspecting, filtering, and possibly blocking suspicious input and output before it reaches the model or the end user. The name deliberately echoes the traditional network firewall: just as the latter controls traffic entering and leaving a network according to predefined rules, a prompt firewall controls the textual traffic entering and leaving an AI system.

Definition

What it is

A prompt firewall is a software layer placed between the user and the language model, dedicated to inspecting, filtering, and possibly blocking suspicious input and output before it reaches the model or the end user. The name deliberately echoes the traditional network firewall: just as the latter controls traffic entering and leaving a network according to predefined rules, a prompt firewall controls the textual traffic entering and leaving an AI system.

How it works

Unlike a single safety filter, a prompt firewall is typically a modular system combining multiple detection techniques: known jailbreak and prompt injection patterns, intent classifiers trained to recognize manipulation attempts, consistency checks between the expected and actually received context, and configurable rules specific to the application domain. It can operate both on user input and on external content an agent retrieves during execution, such as web pages or documents.

Applications

It is particularly relevant for enterprise applications exposing a language model to external users or untrusted content: a prompt firewall reduces the attack surface without requiring changes to the underlying model, which makes it a practical solution even when using third-party models with no direct control. Several providers now offer prompt firewalls as an independent service, deployable in front of any model or API.

History & etymology

The term spread starting in 2023, when the growth of AI agents connected to external tools and data made clear the need for a dedicated, reusable protection layer, distinct from both the model and the application using it.

Definizione (italiano)

Un prompt firewall è un livello software posto tra l'utente e il modello linguistico, dedicato a ispezionare, filtrare ed eventualmente bloccare input e output sospetti prima che raggiungano il modello o l'utente finale. Il nome richiama volutamente il firewall di rete tradizionale: così come quest'ultimo controlla il traffico che entra ed esce da una rete secondo regole predefinite, il prompt firewall controlla il traffico testuale che entra ed esce da un sistema AI.

A differenza di un singolo filtro di sicurezza, un prompt firewall è tipicamente un sistema modulare che combina più tecniche di rilevamento: pattern noti di jailbreak e prompt injection, classificatori di intento addestrati a riconoscere tentativi di manipolazione, controlli di coerenza tra il contesto atteso e quello effettivamente ricevuto, e regole configurabili specifiche per il dominio applicativo. Può operare sia sull'input dell'utente sia sul contenuto esterno che un agente recupera durante l'esecuzione, come pagine web o documenti.

È particolarmente rilevante per le applicazioni aziendali che espongono un modello linguistico a utenti esterni o a contenuti non fidati: un prompt firewall riduce la superficie di attacco senza richiedere di modificare il modello sottostante, il che lo rende una soluzione praticabile anche quando si usano modelli di terze parti su cui non si ha controllo diretto. Diversi fornitori offrono oggi prompt firewall come servizio indipendente, integrabile davanti a qualsiasi modello o API.

Il termine si è diffuso a partire dal 2023, quando la crescita degli agenti AI connessi a strumenti e dati esterni ha reso evidente la necessità di un livello di protezione dedicato e riutilizzabile, distinto sia dal modello sia dall'applicazione che lo utilizza.

Related terms

More in Prompting

Put it into practice

From our network

AGORÀ Intelligence — Enterprise AI Governance Platform

Govern AI at scale: policies, adoption and measurable results on your data. Built for boards and C-suite.

Visit agora-intelligence.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified — the AI coach that trains and certifies your prompt engineering — by Agora Intelligence.