LLM01:2026 Prompt Injection
Any input the model reads (a user message, a retrieved document, a tool result, an image, its own memory) changes its behaviour in a way the developer did not intend, because the model draws no hard line between instructions and data.
- Control
-
Treat every channel into the context as untrusted: input and output guardrails, narrow tool scopes, and a human checkpoint before consequential actions.
- Test
-
- promptfoo
indirect-prompt-injection: plants instructions in prompt variables and checks whether the model follows them - garak
latentinjection: buries injections inside documents such as a resume or a report
- promptfoo
- Evidence for
- Same threat elsewhere
- Layer
- L3 L4
Control-framework ids
- ISO/IEC 42001 Annex A
-
-
A.6.2.4AI system verification and validation -
A.6.2.6AI system operation and monitoring
-
- CSA AICM v1.1 domain
-
-
AISApplication & Interface Security
-
- NIST SP 800-218A task
-
-
PW.1.1Risk modelling (an SSDF 1.1 task; 800-218A recommends including AI-specific threat types)
-
- NIST COSAiS use case
-
- Using a generative AI assistant (LLM)
- AI agent systems: single agent
- Sources
- 1 3 7 8 10 11 13 14