loader
Research
AI Security

Prompt Injection Is an Architecture Problem, Not a Filter Problem

Every proposed fix for prompt injection that operates on the text of a prompt is defeated by rephrasing. The durable mitigations constrain what a model is permitted to do, not what it is allowed to read.

The mistake in the framing

Prompt injection is usually described as a filtering problem: untrusted text reaches a model, the text contains instructions, and the fix is to detect and strip those instructions. Almost every proposed defence follows from that framing — classifiers that score inputs for instruction-like language, delimiters that mark where untrusted content begins, system prompts that tell the model to ignore anything the content says.

All of them share a defect. They assume "an instruction" is a property of text that can be recognised. It is not. An instruction is any text that changes what the model does next, and the space of such text is as large as language itself. A filter tuned on imperative phrasing is defeated by a question. A filter tuned on English is defeated by another language, by base64, by a description of what to do rather than a command to do it, or by text that merely establishes a premise the model then acts on.

The deeper issue is architectural. A language model receives one sequence of tokens. Instructions from the developer and content from the world arrive in the same channel, in the same representation, with no structural marker separating authority from data. Asking the model to honour that distinction is asking it to reconstruct information the interface never carried. It will often succeed, which is what makes the failure mode so dangerous: the defence appears to work until an input is phrased in a way that happens to win.

Why delimiters and system prompts underperform

Wrapping untrusted content in markers and instructing the model to treat everything inside as inert is the most common mitigation, and it is worth being precise about why it is weak rather than useless.

It reduces the rate of accidental compliance. Content that incidentally resembles an instruction — a support ticket quoting an earlier email, a document containing example commands — is much less likely to be acted on. That is a real gain and worth having.

It does not survive an adversary, for two reasons. First, if the delimiter is predictable, the content can contain it and close the boundary early. Second, and more fundamentally, the instruction to ignore the content is itself just text in the same channel, competing on equal footing with the content that says otherwise. There is no mechanism that makes one authoritative. The developer's instruction usually wins because it appears earlier and is phrased with more authority, not because the architecture guarantees it.

Treat delimiters as hygiene rather than as a security control. They lower noise; they do not stop an attacker.

The property that actually matters

The useful question is not "can the model be persuaded to say something unintended" but "what happens when it is". A model that produces the wrong sentence is a quality problem. A model that produces the wrong sentence and can send email, query a database, or spend money is a security problem.

That distinction points directly at the mitigation. Injection is only dangerous in proportion to the capability sitting behind the model. Constrain that capability and the same successful injection becomes an inconvenience rather than an incident.

This reframing is unpopular because it does not promise to solve prompt injection. It concedes the model may be manipulated and asks what the blast radius is when it is. That concession is uncomfortable but honest, and it is the only approach that survives an attacker who is willing to iterate on phrasing.

Constraining capability in practice

Give the model tools with narrow contracts, not broad ones. A tool that executes arbitrary SQL is compromised the moment the model is. A tool that fetches an order by identifier, scoped to the authenticated user, is not — the worst outcome of a successful injection is a legitimate lookup the user was already entitled to make. Design tools so that no argument the model can supply produces an action the user could not perform themselves.

Derive authority from the session, never from the conversation. Whose data the model may access must be established by the authenticated session, not by anything appearing in the prompt. If a request can persuade the model to operate as a different user, the injection has become privilege escalation. This sounds obvious and is violated constantly, usually by passing an identifier through the prompt because it was convenient.

Separate reading from acting. A component that summarises untrusted documents should not also hold credentials that let it write. Where a workflow genuinely requires both, split it: one context ingests hostile content and produces a structured, validated result; a second acts on that result without ever seeing the original text. The attacker can control what the first component reads but not what schema the second accepts.

Require confirmation on the basis of consequence, not confidence. Irreversible or outward-facing actions — sending a message, moving money, deleting records, changing permissions — should require human approval regardless of how certain the model appears. Confidence is not evidence, and a successfully injected model is often extremely confident.

Validate outputs against a schema before they reach anything. If a tool call must be one of a known set with typed arguments in permitted ranges, a great deal of injected instruction becomes unrepresentable. The model cannot invoke what the interface does not expose.

Where filtering still earns its place

None of this makes input inspection worthless; it makes it a layer rather than a boundary.

Classifiers are cheap and catch unsophisticated attempts, which is most volume in practice. Logging inputs that score highly gives visibility into whether anyone is probing at all — frequently the first evidence of a targeted attempt. And in narrow domains where legitimate input has predictable shape, filtering can be genuinely tight.

The failure is architectural: treating the classifier as the control and building capability behind it on the assumption it holds. Layer it over constrained capability and it adds value. Rely on it alone and it postpones an incident rather than preventing one.

What to test

Testing prompt injection by trying known payloads measures whether you are vulnerable to yesterday's phrasing. More useful questions:

  • For each tool, what is the worst outcome if the model invokes it with attacker-chosen arguments? If that answer is unacceptable, the tool is too broad regardless of any filtering in front of it.
  • Can any prompt content cause the system to operate as a different user? If yes, that is the finding — the injection technique is incidental.
  • Which actions are irreversible, and which of those can be triggered without human confirmation?
  • If a classifier were removed entirely, what would change? If the answer is "everything", the classifier is load-bearing and the architecture is wrong.

The honest summary

Prompt injection is not close to solved, and defences framed as detecting malicious text will keep losing to rephrasing, because the category they are trying to recognise is not well defined. What is tractable is limiting what a manipulated model can reach. Systems designed on the assumption the model may be turned against them fail in ways that are visible, bounded and recoverable. Systems that assume the model can be kept honest fail quietly and completely, usually at the moment someone is genuinely trying.