Services · AI security

AI security for systems where the model can act

Adding a language model to an application moves the untrusted input surface somewhere your existing controls do not watch. Most teams test the model when they should be testing the tools, the retrieval index and the permissions behind both.

OWASP Top 10 for LLM ApplicationsNIST AI RMF alignedTesting and remediation
Quick answer

AI security is the practice of testing and defending applications that use a language model, where the attack surface is the text the model reads and the tools it can call. The threat model differs from a normal web application: untrusted content becomes instruction, retrieval becomes an injection channel, and an agent's permissions become its blast radius. We test for those failures and build the controls that contain them.

Threat model

What actually changes when you add a model

Four things move, and they move at once.

The untrusted input surface relocates

The model reads a context window assembled from the system prompt, the conversation, retrieved documents, tool results and whatever a previous step wrote down. All of it is input, and most is never validated, because nobody thinks of a search result as input. The perimeter is the context window now.

The model becomes a confused deputy

The model runs with the application's authority, not the requester's. Anyone who can influence the text it reads borrows that authority. This is the confused deputy problem restated in natural language, and it is why scope design matters more here than prompt wording ever will.

Retrieval becomes an injection vector

Retrieval-augmented generation is the standard way to ground answers in your own content, and we build it constantly on the AI consulting side of the firm. It also loads third-party text into the model's instruction channel by design. A supplier PDF, a scraped page, a wiki anyone can edit: each is an execution path that looks like nothing on an architecture diagram.

The agent's tools are the blast radius

A model that can only produce text has a bounded worst case. One that can send email, write to a database or move money has a worst case defined by those permissions combined. When we review an AI agent, the first artefact we ask for is the tool manifest and the service account it runs as. That page usually predicts the findings before testing starts.

Vocabulary

The attack classes, defined

We use the OWASP Top 10 for LLM Applications as the working taxonomy. These are its categories in our own words.

Prompt injection
Instructions from an untrusted source that the model treats as authoritative. Direct injection comes from the user; indirect injection arrives inside content the system retrieves or browses.
Insecure output handling
Model output passed downstream without encoding or validation. Rendered as markup it is cross-site scripting; in a query builder it is injection; in a shell it is remote code execution.
Excessive agency
The system can do more than the task requires: too many tools, scopes wider than necessary, no confirmation before irreversible actions, no ceiling on the actions one request can trigger.
Sensitive information disclosure
The model reveals secrets, personal data or another tenant's content, because it was in context, because retrieval ignored entitlements, or because output carried it somewhere it should not go.
Supply chain and provenance risk
Models, adapters, embeddings, extensions and connector servers trusted implicitly. The dependency graph now includes weights and prompts, which are harder to review than a lock file.
Poisoning and unbounded consumption
Content planted where it will be ingested and shape future answers, and requests engineered to be expensive or to loop. One costs accuracy, the other a bill and an outage.
The main event

Direct and indirect prompt injection

Direct prompt injection is the one everybody has seen. A user types something designed to override the application's instructions and the assistant starts treating them as its author. It is the less dangerous half, because the attacker is attacking their own session.

Indirect prompt injection is the class that changes how you design. The hostile instructions are not typed by anyone. They sit inside content the system fetches on its own initiative.

A concrete example, described rather than demonstrated

An internal assistant answers questions over a document library. Suppliers upload tenders to a shared folder that feeds the index. One document contains, in white text at eight point, a paragraph addressed to the model rather than the reader: a claim that earlier instructions are obsolete, then a request to attach a named internal file to any summary and send it onward. An analyst asks the assistant to compare the tenders. It retrieves the poisoned document, reads that paragraph as instruction, and attempts the action with the analyst's session and the application's tool permissions.

The same shape applies to an agent that browses. It researches a competitor and visits a page containing text written for the agent rather than the reader. Nothing was typed by a user. Nothing crossed your login. The attacker only had to publish.

This cannot be filtered reliably because there is no formal grammar separating instruction from information in English. A classifier that flags imperative text flags your own documents, and a paraphrase gets past it. Detection is one layer, and we deploy it. It is not the layer.

So the design question is not whether an injection can succeed. Assume it will. The question is what the model can do in the instant it is convinced, and how fast you would know. Both are questions about tools, scopes, egress and logging.

Reference

Attack, symptom, control

The right-hand column is deliberately unexciting. Nearly every control is an access control, an allowlist or an encoding step.

Attack classWhat it looks like in practiceThe control that contains it
Direct prompt injectionA user talks the assistant past its brief, into ignoring policy or revealing its configurationSecrets out of context, policy enforced in code outside the model, system prompt treated as public
Indirect prompt injectionA poisoned document or web page issues instructions the assistant follows while summarising itRetrieved content marked as data, hidden text stripped, confirmation before any tool call it triggers
Data exfiltration through outputSensitive content steered into a link, an image reference or an outbound messageEgress allowlist, no automatic fetching of model-supplied URLs or images, outbound inspection
Excessive agencyAn agent with a delete, payment or messaging tool acts irreversibly on a bad instructionTool allowlists, narrow scopes, human confirmation, hard action and spend caps
Insecure output handlingOutput rendered as markup or passed into a query or a shell, producing injection downstreamEncode at every sink, never execute model text, parameterise queries, validate against a schema
Over-permissioned retrievalA junior employee gets an answer sourced from a board pack, because the index ignores permissionsPermission-aware retrieval evaluated at query time, tested against a matrix of user roles
Supply chain compromiseA model, adapter, extension or connector server behaves differently from its descriptionPinned versions, provenance checks, an internal registry, review of every tool an agent can reach
Unbounded consumptionTraffic or a looping agent drives token spend until the service degradesRate limits, token budgets, recursion caps, alerting on cost per session

Categories align to the OWASP Top 10 for LLM Applications.

Defence

The boring controls that do the work

None of this is novel, and that is the point. These are controls a competent platform team already knows how to build, applied where nobody applied them.

Engagement

What a red team engagement looks like

Two to four weeks for one application, longer where several agents share a tool layer.

  1. 01

    Map the system before touching it

    Trust boundaries, identities, tool manifest, data sources, retrieval design and where output is rendered. Half the findings are visible on this diagram, so we draw it with your engineers.

  2. 02

    Agree rules of engagement in writing

    Environment, dataset, which tests are destructive, who gets called if something breaks, and what counts as a finding.

  3. 03

    Test the classes, not the folklore

    Injection through every ingestion path, exfiltration through output, tool abuse and scope escalation, retrieval permissions across user roles, output handling, supply chain and cost.

  4. 04

    Report so an engineer can act

    Each finding carries its trigger conditions, observed impact, a severity reflecting your context, and the control that removes it. Attacks are described at the level of class and effect. We do not publish working payloads.

  5. 05

    Remediate and retest

    We implement the fixes or review yours, rerun the suite, and hand it over so it runs on every release. A one-off test on a system whose model changes monthly is a snapshot of a moving thing.

Darren consistently demonstrated exceptional technical skill and a strong understanding of various jurisdictions' legal requirements.

VicVP and Chief Data Privacy Officer, Circana
Scope

What we test, and what we fix

Most engagements use three or four of these, decided in the mapping week.

LLM application red teaming

Assistants, copilots and internal tools. Injection through every ingestion path, exfiltration through output, refusal behaviour verified rather than assumed.

Retrieval and permission review

How the corpus is built, who can write to it, and whether retrieval enforces entitlements at query time.

Agent and tool layer assessment

Tool manifests, scopes, confirmation flows, loop limits, audit trails, and what the agent's service account can reach on a bad day.

AI supply chain assessment

Model and adapter provenance, third-party extensions and connector servers, and the review process for anything new an agent may call.

Guardrail and control engineering

We build remediation as often as we report it: permission-aware retrieval, allowlists, output encoding, confirmation steps, egress control, budgets and logging.

Context

Where this practice comes from

Green Arrow Consultancy has been responsible for other organisations' web estates since 2012, and by 2022 the portfolio had passed two hundred managed websites. In 2024 the firm partnered with the UpGuard security platform to extend that protection for enterprise clients. AI security sits on top of that decade, because a large share of AI findings are not AI findings at all. They are a service account with too many rights, an index built from a folder anyone can write to, or a rendering layer that trusts whatever it is handed.

Findings are only useful when somebody owns the risk, so every engagement maps into the client's AI governance programme and, where relevant, into the NIST AI Risk Management Framework or an ISO/IEC 42001 management system. Security proves controls exist. Governance decides who is accountable when they fail.

Questions

Frequently asked questions

More on scoping, pricing and ways of working is in the full FAQ.

What is AI security, and how is it different from application security?

AI security covers the failure modes that appear once a language model sits inside an application. Ordinary application security assumes a clear line between code and data. A language model erases it, because instructions and data arrive in the same channel. What is added is this: content the model reads changes what the model does, and its tools decide how far that goes.

What is prompt injection?

Prompt injection is any technique that gets a model to follow instructions from someone other than the application owner. Direct injection comes from the person typing into the interface. Indirect injection arrives inside content the system pulls in by itself: a document in the corpus, a web page an agent browses, an email the assistant can read. Indirect is the more serious of the two, because the attacker never has to touch your product.

Can prompt injection be solved completely?

No, and any supplier who says otherwise is selling something. No filter reliably separates instruction from data in natural language, because the ambiguity is a property of language rather than a bug in the model. What you can do is make a successful injection worthless: secrets out of context, the narrowest set of tools, human confirmation on anything irreversible, restricted egress, and a log of every tool call. Containment is the strategy.

What does an AI red team engagement involve?

Threat modelling to map trust boundaries, tools, data sources and identities. Then testing against agreed rules of engagement: direct and indirect injection, exfiltration through output, tool abuse, over-permissioned retrieval, insecure output handling and unbounded consumption. You get findings ranked by impact, a fix list split into code and configuration changes, and a retest.

Which frameworks do you test against?

The OWASP Top 10 for LLM Applications is the working checklist for technical findings, and the closest thing the field has to a shared vocabulary. We map findings up to the NIST AI Risk Management Framework so risk owners read them in the same language as the rest of the programme, and to ISO/IEC 42001 where a client is building a certifiable AI management system. Web and cloud testing standards apply underneath, because most AI applications are still web applications.

Do you test agents that take actions, not just chatbots?

Yes, and those matter most. A read-only assistant tricked into saying something wrong is an embarrassment. An agent with a payment tool, a delete tool or a mailbox is an incident. We test the tool layer specifically: what the agent can call, what scopes those calls carry, whether a confirmation step can be talked past, and what the audit trail captures.

How often should AI systems be retested?

At least annually, and on any change that widens the blast radius: a new tool, a new data source, a new integration, or a change to who can use the system. Model updates are the one people forget. A vendor improving a model can change refusal behaviour in ways that invalidate a guardrail you tested six months ago.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Find out what your AI system can be talked into

Send us the tool manifest and an architecture sketch. We will tell you where the blast radius is and what a first engagement would cover.