Adding a language model to an application moves the untrusted input surface somewhere your existing controls do not watch. Most teams test the model when they should be testing the tools, the retrieval index and the permissions behind both.
AI security is the practice of testing and defending applications that use a language model, where the attack surface is the text the model reads and the tools it can call. The threat model differs from a normal web application: untrusted content becomes instruction, retrieval becomes an injection channel, and an agent's permissions become its blast radius. We test for those failures and build the controls that contain them.
Four things move, and they move at once.
The model reads a context window assembled from the system prompt, the conversation, retrieved documents, tool results and whatever a previous step wrote down. All of it is input, and most is never validated, because nobody thinks of a search result as input. The perimeter is the context window now.
The model runs with the application's authority, not the requester's. Anyone who can influence the text it reads borrows that authority. This is the confused deputy problem restated in natural language, and it is why scope design matters more here than prompt wording ever will.
Retrieval-augmented generation is the standard way to ground answers in your own content, and we build it constantly on the AI consulting side of the firm. It also loads third-party text into the model's instruction channel by design. A supplier PDF, a scraped page, a wiki anyone can edit: each is an execution path that looks like nothing on an architecture diagram.
A model that can only produce text has a bounded worst case. One that can send email, write to a database or move money has a worst case defined by those permissions combined. When we review an AI agent, the first artefact we ask for is the tool manifest and the service account it runs as. That page usually predicts the findings before testing starts.
We use the OWASP Top 10 for LLM Applications as the working taxonomy. These are its categories in our own words.
Direct prompt injection is the one everybody has seen. A user types something designed to override the application's instructions and the assistant starts treating them as its author. It is the less dangerous half, because the attacker is attacking their own session.
Indirect prompt injection is the class that changes how you design. The hostile instructions are not typed by anyone. They sit inside content the system fetches on its own initiative.
An internal assistant answers questions over a document library. Suppliers upload tenders to a shared folder that feeds the index. One document contains, in white text at eight point, a paragraph addressed to the model rather than the reader: a claim that earlier instructions are obsolete, then a request to attach a named internal file to any summary and send it onward. An analyst asks the assistant to compare the tenders. It retrieves the poisoned document, reads that paragraph as instruction, and attempts the action with the analyst's session and the application's tool permissions.
The same shape applies to an agent that browses. It researches a competitor and visits a page containing text written for the agent rather than the reader. Nothing was typed by a user. Nothing crossed your login. The attacker only had to publish.
This cannot be filtered reliably because there is no formal grammar separating instruction from information in English. A classifier that flags imperative text flags your own documents, and a paraphrase gets past it. Detection is one layer, and we deploy it. It is not the layer.
So the design question is not whether an injection can succeed. Assume it will. The question is what the model can do in the instant it is convinced, and how fast you would know. Both are questions about tools, scopes, egress and logging.
The right-hand column is deliberately unexciting. Nearly every control is an access control, an allowlist or an encoding step.
| Attack class | What it looks like in practice | The control that contains it |
|---|---|---|
| Direct prompt injection | A user talks the assistant past its brief, into ignoring policy or revealing its configuration | Secrets out of context, policy enforced in code outside the model, system prompt treated as public |
| Indirect prompt injection | A poisoned document or web page issues instructions the assistant follows while summarising it | Retrieved content marked as data, hidden text stripped, confirmation before any tool call it triggers |
| Data exfiltration through output | Sensitive content steered into a link, an image reference or an outbound message | Egress allowlist, no automatic fetching of model-supplied URLs or images, outbound inspection |
| Excessive agency | An agent with a delete, payment or messaging tool acts irreversibly on a bad instruction | Tool allowlists, narrow scopes, human confirmation, hard action and spend caps |
| Insecure output handling | Output rendered as markup or passed into a query or a shell, producing injection downstream | Encode at every sink, never execute model text, parameterise queries, validate against a schema |
| Over-permissioned retrieval | A junior employee gets an answer sourced from a board pack, because the index ignores permissions | Permission-aware retrieval evaluated at query time, tested against a matrix of user roles |
| Supply chain compromise | A model, adapter, extension or connector server behaves differently from its description | Pinned versions, provenance checks, an internal registry, review of every tool an agent can reach |
| Unbounded consumption | Traffic or a looping agent drives token spend until the service degrades | Rate limits, token budgets, recursion caps, alerting on cost per session |
Categories align to the OWASP Top 10 for LLM Applications.
None of this is novel, and that is the point. These are controls a competent platform team already knows how to build, applied where nobody applied them.
Two to four weeks for one application, longer where several agents share a tool layer.
Trust boundaries, identities, tool manifest, data sources, retrieval design and where output is rendered. Half the findings are visible on this diagram, so we draw it with your engineers.
Environment, dataset, which tests are destructive, who gets called if something breaks, and what counts as a finding.
Injection through every ingestion path, exfiltration through output, tool abuse and scope escalation, retrieval permissions across user roles, output handling, supply chain and cost.
Each finding carries its trigger conditions, observed impact, a severity reflecting your context, and the control that removes it. Attacks are described at the level of class and effect. We do not publish working payloads.
We implement the fixes or review yours, rerun the suite, and hand it over so it runs on every release. A one-off test on a system whose model changes monthly is a snapshot of a moving thing.
Darren consistently demonstrated exceptional technical skill and a strong understanding of various jurisdictions' legal requirements.
Most engagements use three or four of these, decided in the mapping week.
Assistants, copilots and internal tools. Injection through every ingestion path, exfiltration through output, refusal behaviour verified rather than assumed.
How the corpus is built, who can write to it, and whether retrieval enforces entitlements at query time.
Tool manifests, scopes, confirmation flows, loop limits, audit trails, and what the agent's service account can reach on a bad day.
Model and adapter provenance, third-party extensions and connector servers, and the review process for anything new an agent may call.
We build remediation as often as we report it: permission-aware retrieval, allowlists, output encoding, confirmation steps, egress control, budgets and logging.
Green Arrow Consultancy has been responsible for other organisations' web estates since 2012, and by 2022 the portfolio had passed two hundred managed websites. In 2024 the firm partnered with the UpGuard security platform to extend that protection for enterprise clients. AI security sits on top of that decade, because a large share of AI findings are not AI findings at all. They are a service account with too many rights, an index built from a folder anyone can write to, or a rendering layer that trusts whatever it is handed.
Findings are only useful when somebody owns the risk, so every engagement maps into the client's AI governance programme and, where relevant, into the NIST AI Risk Management Framework or an ISO/IEC 42001 management system. Security proves controls exist. Governance decides who is accountable when they fail.
More on scoping, pricing and ways of working is in the full FAQ.
AI security covers the failure modes that appear once a language model sits inside an application. Ordinary application security assumes a clear line between code and data. A language model erases it, because instructions and data arrive in the same channel. What is added is this: content the model reads changes what the model does, and its tools decide how far that goes.
Prompt injection is any technique that gets a model to follow instructions from someone other than the application owner. Direct injection comes from the person typing into the interface. Indirect injection arrives inside content the system pulls in by itself: a document in the corpus, a web page an agent browses, an email the assistant can read. Indirect is the more serious of the two, because the attacker never has to touch your product.
No, and any supplier who says otherwise is selling something. No filter reliably separates instruction from data in natural language, because the ambiguity is a property of language rather than a bug in the model. What you can do is make a successful injection worthless: secrets out of context, the narrowest set of tools, human confirmation on anything irreversible, restricted egress, and a log of every tool call. Containment is the strategy.
Threat modelling to map trust boundaries, tools, data sources and identities. Then testing against agreed rules of engagement: direct and indirect injection, exfiltration through output, tool abuse, over-permissioned retrieval, insecure output handling and unbounded consumption. You get findings ranked by impact, a fix list split into code and configuration changes, and a retest.
The OWASP Top 10 for LLM Applications is the working checklist for technical findings, and the closest thing the field has to a shared vocabulary. We map findings up to the NIST AI Risk Management Framework so risk owners read them in the same language as the rest of the programme, and to ISO/IEC 42001 where a client is building a certifiable AI management system. Web and cloud testing standards apply underneath, because most AI applications are still web applications.
Yes, and those matter most. A read-only assistant tricked into saying something wrong is an embarrassment. An agent with a payment tool, a delete tool or a mailbox is an incident. We test the tool layer specifically: what the agent can call, what scopes those calls carry, whether a confirmation step can be talked past, and what the audit trail captures.
At least annually, and on any change that widens the blast radius: a new tool, a new data source, a new integration, or a change to who can use the system. Model updates are the one people forget. A vendor improving a model can change refusal behaviour in ways that invalidate a guardrail you tested six months ago.
Send us the tool manifest and an architecture sketch. We will tell you where the blast radius is and what a first engagement would cover.