AI Dictionary › Prompting

Jailbreak

A jailbreak is an attempt to bypass an AI model's safety protections to make it produce content it would refuse. Techniques range from role-play tricks to sophisticated crafted prompts.

Definition

It differs from prompt injection in its target: jailbreaks attack the model's own rules, injection attacks the application built on the model. For companies the issue is concrete, defenses include robust models, application guardrails, red teaming before public release and continuous monitoring.

Related terms

More in Prompting

Put it into practice

From our network

HSE Genius: AI for Safety Data Sheets

Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.

Visit hsegenius.com →

From the Agora Intelligence blog

More on agora-intelligence.com →

📱 Download the Android app (beta) iOS coming soon

Say what you mean. Get what you need.

Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.