# AI Security and LLM Red Teaming | Green Arrow Consultancy

> AI security testing and defence for LLM applications and AI agents: prompt injection, data exfiltration, excessive agency and over-permissioned retrieval.

Source: https://greenarrow.app/services/ai-security/
Last updated: 2026-09-04
Publisher: Green Arrow Consultancy Ltd, Cardiff, Wales, United Kingdom

---

Services · AI security

# AI security for systems where the model can act

Adding a language model to an application moves the untrusted input surface somewhere your existing controls do not watch. Most teams test the model when they should be testing the tools, the retrieval index and the permissions behind both.

OWASP Top 10 for LLM Applications NIST AI RMF aligned Testing and remediation 
 Book a red team engagement → See the governance side 
 
 
 
 User prompt Retrieved docs Web page 
 
 context window · all of it is input

Tool call requested: send_email

Blocked: not on the allowlist for this agent logged

Quick answer

AI security is the practice of testing and defending applications that use a language model, where the attack surface is the text the model reads and the tools it can call. The threat model differs from a normal web application: untrusted content becomes instruction, retrieval becomes an injection channel, and an agent's permissions become its blast radius. We test for those failures and build the controls that contain them.

## Key points

- The model is rarely the vulnerability. The tools it can call and the documents it may read are.

- Indirect prompt injection, hidden instructions inside content the system fetches by itself, is the class most teams never test for.

- Injection cannot be filtered away, so the strategy is containment: least privilege, allowlists and confirmation on anything irreversible.

- Over-permissioned retrieval is our most common finding, and it is an access control bug rather than an AI one.

- Model output is untrusted input to whatever renders or executes it, which makes insecure output handling a classic web vulnerability with a new source.

## On this page

- What actually changes when you add a model

- The attack classes, defined

- Direct and indirect prompt injection

- Attack, symptom, control

- The boring controls that do the work

- What a red team engagement looks like

- What we test, and what we fix

- Frequently asked questions

Threat model

## What actually changes when you add a model

Four things move, and they move at once.

### The untrusted input surface relocates

The model reads a context window assembled from the system prompt, the conversation, retrieved documents, tool results and whatever a previous step wrote down. All of it is input, and most is never validated, because nobody thinks of a search result as input. The perimeter is the context window now.

### The model becomes a confused deputy

The model runs with the application's authority, not the requester's. Anyone who can influence the text it reads borrows that authority. This is the confused deputy problem restated in natural language, and it is why scope design matters more here than prompt wording ever will.

### Retrieval becomes an injection vector

Retrieval-augmented generation is the standard way to ground answers in your own content, and we build it constantly on the AI consulting side of the firm. It also loads third-party text into the model's instruction channel by design. A supplier PDF, a scraped page, a wiki anyone can edit: each is an execution path that looks like nothing on an architecture diagram.

### The agent's tools are the blast radius

A model that can only produce text has a bounded worst case. One that can send email, write to a database or move money has a worst case defined by those permissions combined. When we review an AI agent, the first artefact we ask for is the tool manifest and the service account it runs as. That page usually predicts the findings before testing starts.

Vocabulary

## The attack classes, defined

We use the OWASP Top 10 for LLM Applications as the working taxonomy. These are its categories in our own words.

**Prompt injection**

Instructions from an untrusted source that the model treats as authoritative. Direct injection comes from the user; indirect injection arrives inside content the system retrieves or browses.

**Insecure output handling**

Model output passed downstream without encoding or validation. Rendered as markup it is cross-site scripting; in a query builder it is injection; in a shell it is remote code execution.

**Excessive agency**

The system can do more than the task requires: too many tools, scopes wider than necessary, no confirmation before irreversible actions, no ceiling on the actions one request can trigger.

**Sensitive information disclosure**

The model reveals secrets, personal data or another tenant's content, because it was in context, because retrieval ignored entitlements, or because output carried it somewhere it should not go.

**Supply chain and provenance risk**

Models, adapters, embeddings, extensions and connector servers trusted implicitly. The dependency graph now includes weights and prompts, which are harder to review than a lock file.

**Poisoning and unbounded consumption**

Content planted where it will be ingested and shape future answers, and requests engineered to be expensive or to loop. One costs accuracy, the other a bill and an outage.

The main event

## Direct and indirect prompt injection

Direct prompt injection is the one everybody has seen. A user types something designed to override the application's instructions and the assistant starts treating them as its author. It is the less dangerous half, because the attacker is attacking their own session.

Indirect prompt injection is the class that changes how you design. The hostile instructions are not typed by anyone. They sit inside content the system fetches on its own initiative.

### A concrete example, described rather than demonstrated

An internal assistant answers questions over a document library. Suppliers upload tenders to a shared folder that feeds the index. One document contains, in white text at eight point, a paragraph addressed to the model rather than the reader: a claim that earlier instructions are obsolete, then a request to attach a named internal file to any summary and send it onward. An analyst asks the assistant to compare the tenders. It retrieves the poisoned document, reads that paragraph as instruction, and attempts the action with the analyst's session and the application's tool permissions.

The same shape applies to an agent that browses. It researches a competitor and visits a page containing text written for the agent rather than the reader. Nothing was typed by a user. Nothing crossed your login. The attacker only had to publish.

This cannot be filtered reliably because there is no formal grammar separating instruction from information in English. A classifier that flags imperative text flags your own documents, and a paraphrase gets past it. Detection is one layer, and we deploy it. It is not the layer.

So the design question is not whether an injection can succeed. Assume it will. The question is what the model can do in the instant it is convinced, and how fast you would know. Both are questions about tools, scopes, egress and logging.

Reference

## Attack, symptom, control

The right-hand column is deliberately unexciting. Nearly every control is an access control, an allowlist or an encoding step.

| Attack class | What it looks like in practice | The control that contains it |
|---|---|---|
| Direct prompt injection | A user talks the assistant past its brief, into ignoring policy or revealing its configuration | Secrets out of context, policy enforced in code outside the model, system prompt treated as public |
| Indirect prompt injection | A poisoned document or web page issues instructions the assistant follows while summarising it | Retrieved content marked as data, hidden text stripped, confirmation before any tool call it triggers |
| Data exfiltration through output | Sensitive content steered into a link, an image reference or an outbound message | Egress allowlist, no automatic fetching of model-supplied URLs or images, outbound inspection |
| Excessive agency | An agent with a delete, payment or messaging tool acts irreversibly on a bad instruction | Tool allowlists, narrow scopes, human confirmation, hard action and spend caps |
| Insecure output handling | Output rendered as markup or passed into a query or a shell, producing injection downstream | Encode at every sink, never execute model text, parameterise queries, validate against a schema |
| Over-permissioned retrieval | A junior employee gets an answer sourced from a board pack, because the index ignores permissions | Permission-aware retrieval evaluated at query time, tested against a matrix of user roles |
| Supply chain compromise | A model, adapter, extension or connector server behaves differently from its description | Pinned versions, provenance checks, an internal registry, review of every tool an agent can reach |
| Unbounded consumption | Traffic or a looping agent drives token spend until the service degrades | Rate limits, token budgets, recursion caps, alerting on cost per session |

Categories align to the OWASP Top 10 for LLM Applications.

Defence

## The boring controls that do the work

None of this is novel, and that is the point. These are controls a competent platform team already knows how to build, applied where nobody applied them.

- Least-privilege retrieval. The index enforces the requester's entitlements at query time, not a post-filter on the answer, which leaks through summaries.

- An explicit tool allowlist. Every tool enumerated, scoped and justified. A tool that exists because it might be useful later does not exist yet.

- Human confirmation on irreversible actions. Sending, paying, deleting and publishing stop for a person, described in plain language rather than raw arguments.

- Output encoding at every sink. Escaped for HTML, parameterised for SQL, never passed to a shell, schema-validated when structured.

- A replayable log of every tool call. Arguments, identity, result and decision. Most AI incidents are unprovable rather than undetectable.

- Rate limits and cost ceilings per identity. Token budgets, request quotas, agent step limits, and an alert when a session costs far more than the median.

- Security tests in the release pipeline. The red team suite runs beside the quality evaluation set on every change and every model upgrade.

Engagement

## What a red team engagement looks like

Two to four weeks for one application, longer where several agents share a tool layer.

- 01

### Map the system before touching it

Trust boundaries, identities, tool manifest, data sources, retrieval design and where output is rendered. Half the findings are visible on this diagram, so we draw it with your engineers.

- 02

### Agree rules of engagement in writing

Environment, dataset, which tests are destructive, who gets called if something breaks, and what counts as a finding.

- 03

### Test the classes, not the folklore

Injection through every ingestion path, exfiltration through output, tool abuse and scope escalation, retrieval permissions across user roles, output handling, supply chain and cost.

- 04

### Report so an engineer can act

Each finding carries its trigger conditions, observed impact, a severity reflecting your context, and the control that removes it. Attacks are described at the level of class and effect. We do not publish working payloads.

- 05

### Remediate and retest

We implement the fixes or review yours, rerun the suite, and hand it over so it runs on every release. A one-off test on a system whose model changes monthly is a snapshot of a moving thing.

> Darren consistently demonstrated exceptional technical skill and a strong understanding of various jurisdictions' legal requirements.

, Vic VP and Chief Data Privacy Officer, Circana

Scope

## What we test, and what we fix

Most engagements use three or four of these, decided in the mapping week.

### LLM application red teaming

Assistants, copilots and internal tools. Injection through every ingestion path, exfiltration through output, refusal behaviour verified rather than assumed.

### Retrieval and permission review

How the corpus is built, who can write to it, and whether retrieval enforces entitlements at query time.

### Agent and tool layer assessment

Tool manifests, scopes, confirmation flows, loop limits, audit trails, and what the agent's service account can reach on a bad day.

### AI supply chain assessment

Model and adapter provenance, third-party extensions and connector servers, and the review process for anything new an agent may call.

### Guardrail and control engineering

We build remediation as often as we report it: permission-aware retrieval, allowlists, output encoding, confirmation steps, egress control, budgets and logging.

Context

## Where this practice comes from

Green Arrow Consultancy has been responsible for other organisations' web estates since 2012, and by 2022 the portfolio had passed two hundred managed websites. In 2024 the firm partnered with the UpGuard security platform to extend that protection for enterprise clients. AI security sits on top of that decade, because a large share of AI findings are not AI findings at all. They are a service account with too many rights, an index built from a folder anyone can write to, or a rendering layer that trusts whatever it is handed.

Findings are only useful when somebody owns the risk, so every engagement maps into the client's AI governance programme and, where relevant, into the NIST AI Risk Management Framework or an ISO/IEC 42001 management system. Security proves controls exist. Governance decides who is accountable when they fail.

Questions

## Frequently asked questions

More on scoping, pricing and ways of working is in the full FAQ.

### What is AI security, and how is it different from application security?

AI security covers the failure modes that appear once a language model sits inside an application. Ordinary application security assumes a clear line between code and data. A language model erases it, because instructions and data arrive in the same channel. What is added is this: content the model reads changes what the model does, and its tools decide how far that goes.

### What is prompt injection?

Prompt injection is any technique that gets a model to follow instructions from someone other than the application owner. Direct injection comes from the person typing into the interface. Indirect injection arrives inside content the system pulls in by itself: a document in the corpus, a web page an agent browses, an email the assistant can read. Indirect is the more serious of the two, because the attacker never has to touch your product.

### Can prompt injection be solved completely?

No, and any supplier who says otherwise is selling something. No filter reliably separates instruction from data in natural language, because the ambiguity is a property of language rather than a bug in the model. What you can do is make a successful injection worthless: secrets out of context, the narrowest set of tools, human confirmation on anything irreversible, restricted egress, and a log of every tool call. Containment is the strategy.

### What does an AI red team engagement involve?

Threat modelling to map trust boundaries, tools, data sources and identities. Then testing against agreed rules of engagement: direct and indirect injection, exfiltration through output, tool abuse, over-permissioned retrieval, insecure output handling and unbounded consumption. You get findings ranked by impact, a fix list split into code and configuration changes, and a retest.

### Which frameworks do you test against?

The OWASP Top 10 for LLM Applications is the working checklist for technical findings, and the closest thing the field has to a shared vocabulary. We map findings up to the NIST AI Risk Management Framework so risk owners read them in the same language as the rest of the programme, and to ISO/IEC 42001 where a client is building a certifiable AI management system. Web and cloud testing standards apply underneath, because most AI applications are still web applications.

### Do you test agents that take actions, not just chatbots?

Yes, and those matter most. A read-only assistant tricked into saying something wrong is an embarrassment. An agent with a payment tool, a delete tool or a mailbox is an incident. We test the tool layer specifically: what the agent can call, what scopes those calls carry, whether a confirmation step can be talked past, and what the audit trail captures.

### How often should AI systems be retested?

At least annually, and on any change that widens the blast radius: a new tool, a new data source, a new integration, or a change to who can use the system. Model updates are the one people forget. A vendor improving a model can change refusal behaviour in ways that invalidate a guardrail you tested six months ago.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed 04 September 2026.

Keep reading

## Related

### AI Governance & Compliance

The layer above this one: inventory, risk classification, oversight and who is accountable.

Read this →

### AI Agents & Automation

How we build agents whose actions are scoped, logged, confirmable and reversible by design.

Read this →

### Privacy Consulting

Most AI security incidents are data incidents. Lawful basis, retention and data flow start here.

Read this →

## Find out what your AI system can be talked into

Send us the tool manifest and an architecture sketch. We will tell you where the blast radius is and what a first engagement would cover.

Book an AI red team engagement → 
 See the AI practice
