A prompt crafted to make a model disregard its safety instructions entirely. OWASP treats jailbreaking as a form of prompt injection 1; it is tested with red-team suites in the eval gate and contained at runtime by guardrails that do not depend on the model's own refusals.
- Developed in
- ch. 14, The test-type matrix
- Chapters
- ch. 14, Development · ch. 17, Incidents
- Contrast with
- Prompt injection
- Source
- 1 numbered reference, listed below
Commonly confused
Prompt injection versus Jailbreak
- Prompt injection
- An input that alters a model's behaviour or output in ways its designers did not intend.
- Jailbreak
- A prompt crafted to make a model disregard its safety instructions entirely.
The differenceAny input that alters behaviour in unintended ways, direct or hidden in processed content, against inputs aimed at dropping the safety rules.
Why it mattersJailbreak evals test refusals; injection also needs least-privilege tools and isolation of untrusted content.
Where it is used
6 chapters of the Body of Knowledge use the term. Each link opens the first section that does.
- 01 · Definition The limits of the eval gate 1 mention
- 03 · Values & Principles The eight values 1 mention
- 04 · The Stack Layer 03: Evals & Red Teaming as Evidence 1 mention
- 14 · Development Testing and validation 1 mention
- 15 · Deployment Model types and deployment options 1 mention
- 17 · Incidents Root-cause analysis 1 mention
Related terms
Sources
- [1] LLM01:2026 Prompt Injection (OWASP Top 10 for LLM Applications 2026; direct and indirect injection; jailbreaking as the subset of prompt injection that aims to make the model violate its safety protocols; entry text in github.com/GenAI-Security-Project/GenAI-LLM-Top10, 2026/final). OWASP GenAI Security Project. 2026-08-03. https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/ (verified: primary)
Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .