AI Dictionary › Prompting
A jailbreak is an attempt to bypass an AI model's safety protections to make it produce content it would refuse. Techniques range from role-play tricks to sophisticated crafted prompts.
It differs from prompt injection in its target: jailbreaks attack the model's own rules, injection attacks the application built on the model. For companies the issue is concrete, defenses include robust models, application guardrails, red teaming before public release and continuous monitoring.
From our network
HSE Genius: AI for Safety Data Sheets
Extract SDS data, H phrases and ECHA compliance checks in seconds, powered by AI.
Visit hsegenius.com →From the Agora Intelligence blog
📱 Download the Android app (beta) iOS coming soon
Say what you mean. Get what you need.
Grace Certified, the AI coach that trains and certifies your prompt engineering, by Agora Intelligence.