Insights · AI Security

Prompt injection, explained for people who have to sign it off

You do not need to understand transformers to make a good decision about this. You need to understand one property of language models, and then ask four questions about what the system in front of you is allowed to touch. This article is written for the person whose name goes on the approval.

Written for non-engineersAI securityNo attack recipes
Quick answer

Prompt injection is when text a system reads gets treated as an instruction to follow rather than information to process. A language model cannot reliably tell your instructions apart from instructions hidden in a document, an email or a web page it has been asked to read. There is no complete fix, so the discipline is containment: limit what the system can reach, limit what it can do, and make the worst case survivable.

The core idea

The one property you need to understand

Think about how you brief a capable new colleague by email. You write what you want, you attach the material, and you trust them to notice the difference between your instructions and the contents of the attachment. A person makes that distinction using context, judgement and an understanding of who is entitled to give them orders.

A language model does not have any of that. Everything it receives, your carefully written system instructions, the user's question, the three documents your search returned, the email body it was asked to summarise, arrives as one continuous stream of text. The model is extremely good at working out what text is asking for. It is not good at working out whether that text had the standing to ask.

So the rule is simple. Any text the system reads is a potential instruction. Not “might be misinterpreted”, not “in unusual circumstances”. If a paragraph inside a document your assistant is summarising is written in the form of a direction, there is a real chance the assistant treats it as one. That is the whole vulnerability, and everything else in this article follows from it.

It is worth sitting with how unfamiliar that is. In conventional software, code and data are separate by construction. A spreadsheet cell cannot tell the accounting system to do something else. Here the separation is not architectural, it is a matter of the model's judgement in the moment, and judgement can be talked round. The industry calls this class of problem prompt injection, and it sits at the top of the OWASP Top 10 for LLM Applications.

The shape of an attack, at a conceptual level, is unremarkable. Somebody puts wording into a document, a web page, a review or a message that is addressed to the machine rather than to the human reader, telling it to disregard its previous brief and do something else instead. It may be visually hidden, it may be phrased as an urgent administrative note, it may be buried in a long file nobody reads to the end. We are deliberately not publishing a worked example, because a worked example is a working attack, and the concept is what you need in order to make the decision.

Taxonomy

Two kinds, and why one matters more

If you remember one distinction from this article, make it this one.

Direct injection

Someone types something into your interface designed to make the assistant break its brief. They are attacking a system they already have access to, so they gain little they did not have. The typical outcome is an off-brand answer and an embarrassing screenshot. Real, worth defending against, rarely a boardroom matter on its own.

Indirect injection

The instruction arrives inside content the system fetches by itself. A document in the search corpus. An inbound email an assistant triages. A web page an agent visits. A support ticket, a product review, a submitted CV, a supplier invoice. The attacker never touches your product, and the person harmed is whoever asks the next question.

Why indirect is the serious one

Three reasons. It scales, because one poisoned document reaches everyone who queries near it. It escalates, because the assistant acts with the permissions of the person asking, who may be far more privileged than the attacker. And it is patient: content planted today can sit in an index until somebody asks the right question.

Where poisoned content comes from

Anywhere your organisation accepts text from outside. Customer support systems, recruitment inboxes, supplier portals, public web pages, shared drives that contractors can write to, and forwarded documents. The route into the corpus is usually mundane and already approved.

The agent case

An agent that browses the internet is, by design, reading text written by strangers and deciding what to do next. That is the highest-exposure configuration in common use, and it deserves the tightest limits on what the agent is allowed to do with what it finds.

The comforting case

A read-only assistant answering from a small, curated, internally authored corpus, holding no secrets and able to take no action, has a genuinely modest exposure. Not zero, but proportionate. Knowing which case you are in is most of the work.

Proportionality

Risk follows capability, not knowledge

The instinctive question when people first hear about this is “what could it be tricked into saying?” That is the wrong axis. The right question is “what could it be tricked into doing?”

A system that can only read and answer has one realistic failure mode: it says something wrong, off-brand or disclosing. That is a communications problem, and communications problems are recoverable. A system that can send an email, write to a record, place an order, move money, delete a file or call another system has failure modes that are not recoverable in the same way. The model is identical in both cases. The difference is entirely in what it was connected to.

This is the concept the security community calls excessive agency, and it is the reason the same underlying flaw ranges from trivial to severe depending on the deployment. It also explains why the sensible governance question is not “is the model safe”. It is “what is the largest irreversible action available in this system, and who authorised granting it?”

A second axis is identity. Most assistants act with borrowed authority: they run queries and take actions as though they were the person using them. If a director asks a question and a poisoned document steers the assistant, the resulting action carries a director's permissions. The attacker supplied a paragraph. The system supplied the privilege. That asymmetry is what makes indirect injection worth taking seriously even when the attacker is unsophisticated.

The practical consequence for sign-off is that you should be reviewing a capability list, not a model choice. Which tools has this system been given, what is the worst outcome of each one being called by mistake, and which of those outcomes can be undone. That review is described in more detail under AI security, and the documentation it should produce is covered under AI governance.

Risk

What could actually go wrong

Read down the first column until you reach the most capable thing your system can do. That row is your risk.

What the system is allowed to doWorst realistic outcome of a successful injectionThe control that bounds it
Read a curated internal corpus and answerAn off-brand, incorrect or oddly worded answer that somebody screenshotsCurate the corpus, keep secrets out of the context window, and monitor answers
Search across systems using the asker's identityContent is surfaced to someone entitled to it but summarised in a misleading wayPermission-aware retrieval and citation, so every claim is traceable to a source document
Read inbound email, tickets or submitted documentsUntrusted text from outside steers the assistant on behalf of an internal userTreat all inbound content as hostile input, isolate it from privileged context, and never grant it authority
Browse the public web during a taskA page written by a stranger redirects the task, potentially towards data the agent already holdsRestrict which domains can be reached, strip retrieved content of authority, and forbid secrets in the same context
Draft outbound messages for a human to reviewA misleading or damaging draft is sent because the reviewer approved without reading properlyMake the review meaningful: show what changed, show the source, and make approval an explicit act
Send messages or post content on its ownData leaves the organisation, or a customer receives something that should never have been sentRemove the capability, or bound it to a fixed template, a fixed recipient list and a rate limit
Write to records in a business systemRecords are altered quietly and the change is discovered weeks later during a reconciliationWrite to a staging area, require human confirmation, log every write, and keep the change reversible
Spend money, place orders or move fundsA financial transaction executes that nobody intended and that cannot be recalledHuman confirmation for every transaction, hard value limits, and a separate identity from the assistant
Delete data or revoke accessDestruction that is expensive or impossible to reverse, and an availability incidentDo not grant it. Where it is unavoidable, use soft deletion, a delay, and dual approval
Call other internal systems as an administratorLateral movement across the estate with privileges no individual user holdsNever give an assistant a shared administrative identity. Scope it to the narrowest role that works

Worst realistic outcomes, not worst conceivable ones. The point of the table is proportionate approval, not alarm.

Defence

The controls, in plain language

Ten controls. None of them requires you to understand the model, and all of them are decisions a business can make.

Approval

What to ask before you sign

Seven questions. You can ask all of them without knowing anything about machine learning.

  1. 01

    What can this system read, and who wrote it?

    Ask for the source list. Then ask which of those sources can be written to by someone outside the organisation, including customers, suppliers, candidates and contractors. If the answer is “we pointed it at the shared drive”, the real answer is that nobody knows.

  2. 02

    What can it do, and which actions cannot be undone?

    Ask for the tool list in plain language, one line each. Mark the irreversible ones. If any irreversible action can happen without a person confirming it, that is the finding, and it is a design change rather than a tuning exercise.

  3. 03

    Whose permissions does it act with?

    It should have its own narrow identity, or act strictly within the entitlements of the person asking. If it holds a shared account with broad access because that was simpler to set up, the blast radius of any injection is the whole estate.

  4. 04

    What is the worst thing that happens if it follows an instruction we did not write?

    Ask the team to answer this out loud, specifically, with a scenario. A team that has thought about containment answers in thirty seconds. A team that has not will start by explaining why it is unlikely, which is the wrong answer to a question about consequence.

  5. 05

    Show me the log of a real action.

    Not the design for logging. An actual record: tool called, arguments, requesting identity, result, timestamp. If that record does not exist today, you cannot investigate an incident, and you will find that out at the worst possible moment.

  6. 06

    Who is on the other end when it goes wrong?

    A named owner, an escalation route a front-line person can use, and a decision already taken about who can switch the system off and how quickly. Agree the kill switch before you need it, not during.

  7. 07

    When is it tested again?

    Adversarial testing is not a launch gate you pass once. Every new connector, new tool and model change reopens the question. Put a cadence in the approval, and make the retest a condition of the sign-off rather than a good intention.

Candour

The honest state of the art

It would be more comfortable to end with a product recommendation. There is not one. Input filters, output classifiers, instruction hierarchies and specially trained guard models all raise the effort required and all of them are worth deploying, but they are heuristics applied to natural language, and natural language has an unlimited number of ways to express the same intent. Published research keeps demonstrating this, and the defences keep being partial.

That is not a reason to avoid these systems. It is a reason to deploy them the way we already deploy other things that cannot be made perfectly safe. Nobody expects an email gateway to stop every phishing message. We accept a residual rate and we contain it: limited privileges, confirmation on payments, monitoring, and a rehearsed response. The same posture works here, and it is a posture a board already understands.

The practical result is that the interesting question at sign-off is not whether the model can be tricked. Assume it can. The question is what the organisation has arranged to be true when it is. If the honest answer is that somebody gets a strange answer and the monitoring picks it up, approve it. If the honest answer is that money moves or data leaves, send it back, and be specific about which capability has to change.

One further point worth making to any team building in this space. The corpus decisions and permission decisions that bound this risk are the same decisions that determine your privacy exposure, which is why we treat them together. Those are set out in the twelve privacy questions. And if a supplier is proposing to solve a knowledge problem by training the model rather than by retrieving from a controlled corpus, the trade-offs are covered in retrieval or fine-tuning. Broader definitions are in the glossary.

Questions

Frequently asked questions

More on how we test this under AI security.

What is prompt injection, in one paragraph?

A language model reads everything it is given as one continuous stream of text. Your instructions and the content it has been asked to work on arrive in the same channel, and the model has no reliable way to tell which is which. Prompt injection is what happens when text that was supposed to be data behaves like an instruction instead. If a system reads a document, an email or a web page, whatever is written in that material is a candidate instruction, whether you intended it to be or not.

How is this different from a normal software vulnerability?

Conventional vulnerabilities are mistakes. Somebody forgot to check an input length, and when the mistake is fixed the vulnerability is gone. Prompt injection is not a mistake, it is a consequence of how language models work. Following instructions expressed in natural language is the capability you are paying for. The same property that lets a model take direction from you in plain English lets it take direction from a paragraph inside a support ticket. There is no patch for that, which is why the discipline is containment rather than elimination.

What is the difference between direct and indirect injection?

Direct injection is a user typing something into your interface to make the assistant misbehave. It is mostly a brand and content problem: the user is attacking a system they are already allowed to use, and the worst outcome is usually an embarrassing screenshot. Indirect injection arrives inside content the system pulls in on its own, a document in a retrieval corpus, an inbound email, a web page, a support ticket, a product review, a resume. That is the serious one, because the attacker never touches your product and the victim is often someone with more access than the attacker has.

Can it be fixed?

Not completely, and any supplier telling you otherwise is selling something. Filters and classifiers catch obvious attempts and are worth having, but they are heuristics operating on natural language, and natural language has unlimited ways to express the same intent. What can be fixed is the consequence. Keep secrets out of the context, limit what the model can reach, limit what it can do, require a person to confirm anything irreversible, and log everything. Then a successful injection produces a bad sentence rather than a bad transaction.

Which questions should I ask before approving a deployment?

Four. What can this system read, and could any of it have been written by someone outside the organisation? What can it do, and which of those actions cannot be undone? Whose permissions does it act with, and are they wider than the person asking? And what is the worst thing that happens if it follows an instruction we did not write? If nobody can answer the fourth question without leaving the room, the system is not ready for sign-off.

Does this apply to us if we only have a simple chatbot?

It applies, but the stakes are proportionate. A chatbot that answers from a curated set of published policies, holds no secrets, and cannot take any action has a genuinely small blast radius. The worst realistic outcome is that somebody makes it say something silly and posts it. That is a reputational matter, not a breach. Risk climbs sharply the moment the system reads content from outside the organisation, or gains the ability to send, write, buy or delete.

What is excessive agency?

It is the pattern where a system has been given more capability than its task requires, usually because it was easier to grant broad access than to work out the narrow set. An assistant that only needs to read three folders is given the whole drive. One that needs to draft replies is given the ability to send them. Excessive agency is what converts a prompt injection from an odd answer into an incident, and it appears on the OWASP Top 10 for LLM Applications for exactly that reason.

Should model output ever be treated as trusted?

No. Treat everything a model produces the way you would treat text submitted by an anonymous member of the public, because in a successful injection that is effectively what it is. Do not execute it, do not render it as trusted markup, do not pass it into a database query or a shell without the same handling you would apply to any untrusted input, and be careful about output that can cause an outbound request. This is ordinary application security practice, and it is the control most often skipped because the output feels like it came from inside the system.

How do you test for this?

Threat modelling first, to map what the system reads, what it can do, and whose identity it acts with. Then adversarial testing in a non-production environment against agreed rules of engagement, covering both direct and indirect paths, exfiltration routes, tool misuse and over-broad retrieval. The output is a ranked list of reproducible findings with fixes separated into code changes and configuration changes, and a retest. We describe how we run this under AI security.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Find out what your system would actually do

We threat model the deployment, test it adversarially in a non-production environment, and give you a ranked list of findings with fixes. Then we retest.