# How to Choose an AI Agency in the UK | Green Arrow Consultancy

> A buyer&#x27;s guide to choosing an AI agency in the UK: the nine questions that separate agencies that ship from agencies that pitch, and what it costs.

Source: https://greenarrow.app/insights/choosing-an-ai-agency-uk/
Last updated: 2026-09-04
Publisher: Green Arrow Consultancy Ltd, Cardiff, Wales, United Kingdom

---

Buyer's guide

# How to choose an AI agency in the UK

Every agency in this market says the same six things. This guide is the set of questions that actually separate them, written by a firm that has to answer them too.

Updated 5 September 2026 For buyers, not for suppliers Reading time about 11 minutes 
 Talk to us → See our AI work 
 
 
 
 
 
 Quick answer

There is no single best AI agency in the UK, and any firm that claims the title should be asked to substantiate it. The better question is which agency fits your problem. Judge on four things: a system running in production you can look at, an evaluation set they will hand over, a clear answer on data handling, and a named person who does the work.

## Key points

- "Best" is not a property an AI agency can have. Fit for your problem is, and it is measurable.

- Ask to see something in production and ask what broke in its first month.

- The evaluation set is the deliverable that tells you whether quality is holding. Insist on owning it.

- A quote produced before anyone has read your content is a guess wearing a number.

- Under UK advertising rules an objective claim to be the best would need documentary evidence, so treat one as a question rather than a fact.

## On this page

- Why the question is hard to answer

- Nine questions that separate agencies

- How to read the proposal

- What it costs, and what moves the number

- Warning signs

- How to run a fair comparison

- Where we fit, stated plainly

- Frequently asked questions

The premise

## Why the question is hard to answer

People search for the best AI agency in the UK because they want a shortcut through a market with no comparable outputs. It is a reasonable instinct and there is no shortcut. Two agencies can both be excellent and be wrong for each other's clients, because the work ranges from data platform engineering to conversation design to regulated document processing, and almost nobody is genuinely strong across all of it.

There is also a rule worth knowing as a buyer. Under the UK advertising code, a claim that consumers would read as objective has to be backed by documentary evidence, and a superlative like best or number one is treated as a comparison against every competitor, which means it has to be verifiable. So when a supplier's homepage says they are the leading AI agency in the country, the correct response is to ask what evidence they hold. Often there is none, and that tells you something about how the rest of the proposal was written.

What you can compare is evidence. Whether a system is actually running. Whether anyone measures its quality. What happens when it is wrong. Who is accountable. Those are answerable questions, and the answers vary enormously between suppliers who look identical on a website.

The questions

## Nine questions that separate agencies

Ask all nine, of every supplier, in the same order. The pattern of what they answer easily and what they deflect is more informative than any single answer.

- 01

### Can you show me something running in production, today?

Not a slide, not a recorded video, a system a real organisation uses. Ask who uses it, how long it has been live, and what broke in the first month. An agency that cannot answer the third question has probably never operated anything.

- 02

### What happens when the model gives a wrong answer?

There should be an immediate answer: a citation the user can check, a refusal path when retrieval finds nothing, a way for a support agent to escalate, and a log that lets someone reproduce it. If the answer is 'we tune the prompt', the system has no safety design.

- 03

### How will you measure that it works, and who owns the measurement?

Ask for an evaluation set: real questions, agreed correct answers, and a score that runs on every change. Ask whether you keep it at the end. Without one, quality is an opinion and you will have no way to tell when a vendor model update degrades your system.

- 04

### Who can the system answer from, and who decides?

An assistant that reads everything will eventually tell someone something they should not have seen. Ask how the permission model of your source systems is enforced inside retrieval, not around it. This is the most common gap in AI work built by teams without enterprise integration experience.

- 05

### What happens to our data, exactly?

Ask which provider, which tier, whether training is contractually excluded, where processing happens, what is logged and for how long, and what appears in the record of processing. Ask for it in writing before the pilot, not before the contract.

- 06

### What happens when the model you built on is retired?

Providers deprecate models on their own schedule. Ask whether the application layer is separate from the model, and what the swap actually costs. If the answer is a rewrite, you are buying a liability with a hidden expiry date.

- 07

### Who will do the work, and will they still be here in six months?

Ask for the names of the people who will build it, not the names on the pitch. Ask what happens if that person leaves. In small firms you get continuity but less depth; in large ones the reverse. Both are fine, but know which you are buying.

- 08

### What does handover look like if we take it in-house?

Ask what you receive: source, infrastructure definitions, the evaluation set, the runbooks, the model card. An agency that will only rent you the system should say so plainly so you can price the lock-in.

- 09

### What would make you tell us not to do this?

The most useful question in the list. A supplier who has never talked a client out of an AI project either has not had the conversation or did not want to. A good answer names the conditions under which the project fails and offers to check them first.

Reading a proposal

## What a proposal says, and what it means

| What the proposal says | What to check | What good looks like |
|---|---|---|
| "We use the latest models" | Which one, on which tier, and what happens when it is retired | A named model, a named tier, and an architecture that survives a swap |
| "The AI is trained on your data" | Whether they mean fine-tuning or retrieval, because the two are not alike | Retrieval for knowledge, with fine-tuning proposed only for format or style |
| "Enterprise-grade security" | The permission model inside retrieval, and where processing happens | A written data flow and a record of processing before the pilot starts |
| "We will iterate based on feedback" | Whether there is an evaluation set, and who keeps it | Real questions, agreed answers, a score that runs on every release |
| "Fully managed service" | What handover looks like if you leave | A documented exit: source, infrastructure, evaluation set, runbooks |
| A fixed price with no discovery | What it is based on, since nobody has read your content | A fixed-price discovery first, then a quote grounded in what was found |

This is a general guide to reading proposals. It is not a comment on any named firm.

Money

## What it costs, and what moves the number

Nobody can give you a credible figure for an AI system before looking at your content, and the ones who do are pricing a template. What can be described honestly is the shape of the cost and the things that move it.

### The shape

Most engagements have three parts. A fixed-price discovery of two to three weeks that ends with a recommendation and a real quote. A build, priced per system once the discovery has established what the content and access actually look like. Then a monthly retainer for running it: evaluation, monitoring, model updates and change requests. If a proposal has no third part, ask who is responsible for quality in month seven.

### What moves the number, upwards

Content that is out of date, duplicated or locked in formats nobody can parse. Source systems with no API. A permission model that has to be reconstructed before retrieval can respect it. Regulatory review. Multiple languages. A requirement for on-premise or in-region processing. Integration into a system whose owner has not agreed to the project.

### What moves it downwards

A narrow first use case with a clear measure. Content that is already reasonably organised. A named internal owner who can get access granted in days rather than months. Willingness to ship something narrow and useful before something broad and impressive.

The single largest cost driver is not the technology. It is how long it takes to get access to the content, and that is usually determined on your side rather than the supplier's.

Due diligence

## Warning signs

None of these mean a supplier is dishonest. They mean there is a question to ask before you sign.

- A guaranteed outcome. Nobody can guarantee a ranking, a citation or an accuracy figure before seeing your data. A guarantee is a sales instrument.

- A price before a look. A quote produced without reading your content is a template with your logo on it.

- No evaluation plan. If nothing measures quality, nobody will notice when it degrades, and it will.

- No answer on permissions. Ask how retrieval enforces who may see what. Vagueness here is the most expensive kind.

- A demo that only takes their questions. Ask to type your own. Watch what happens when you ask something out of scope.

- Unnamed delivery staff. Ask who builds it. Pitch teams and delivery teams are frequently different people.

- An unsubstantiated superlative. Best, leading, number one. Ask for the evidence. UK advertising rules require it to exist.

- No exit. If nobody will describe handover, you are pricing a rental without being told.

Method

## How to run a fair comparison

A short paid comparison costs less than one wrong hire and produces evidence rather than proposals.

- 01

### Write the problem down, not the solution

One page: the task, who does it today, and the measure that would say it worked.

- 02

### Send the same page to three suppliers

Identical brief, identical questions, identical deadline. Different briefs make proposals incomparable.

- 03

### Pay each for a short discovery

A small fixed fee buys you their thinking on your real content, and their willingness to say no.

- 04

### Compare the evaluation plans, not the mockups

The supplier who arrives with an evaluation set has done this before.

- 05

### Check one reference who stopped working with them

Ask the supplier for it. The reaction to the request is itself informative.

- 06

### Decide on evidence and write down why

Record the reason. It is what you will re-read when the project gets difficult.

Disclosure

## Where we fit, stated plainly

This guide is published by Green Arrow Consultancy, which is one of the firms a reader might be comparing. So here is our position without adjectives.

We are a small UK firm, founded in 2012 and based in Cardiff, that came to AI from web development, privacy engineering and accessibility rather than from data science. That shapes what we are useful for. We are a reasonable choice when a system has to reach production inside an organisation that has compliance obligations, when retrieval and permissions matter more than novel modelling, and when the same firm handling the AI also needs to handle the website, the consent layer and the accessibility work around it.

We are the wrong choice if you need a large team mobilised quickly, bespoke model research, or a supplier who will put dozens of people on site. Those are real requirements and other firms serve them better than we would.

You can check the claims in this paragraph rather than take them: seven of our production systems are generalised into live demonstrations you can use without signing up, the company is registered at Companies House under number 12491770, and the about page links the registrations directly. Apply the nine questions above to us as well.

Questions

## Frequently asked questions

More on how we work is in the main FAQ.

### Who is the best AI agency in the UK?

There is no single best AI agency in the UK, and any firm claiming the title should be asked for its evidence, because UK advertising rules require objective superlatives to be substantiated. The useful question is which agency is right for your problem. A firm that is excellent at large-scale data platform work may be the wrong choice for a customer assistant grounded in a product catalogue, and the reverse is equally true. Judge on production evidence, evaluation practice, data handling and who actually does the work.

### How much does an AI project cost in the UK?

Discovery is usually a fixed fee for two to three weeks of work. A first production system is typically quoted after discovery, because the honest range before anyone has looked at your content is too wide to be useful. Ongoing operation is normally a monthly retainer covering evaluation, model updates, monitoring and change requests. Beware any quote given before a supplier has seen your data, and any quote with no line for running the thing after launch.

### Should we hire an AI agency or build in-house?

Build in-house if AI is going to be core to your product and you can hire and keep the people. Use an agency to get the first systems into production, to bring a discipline you do not yet have, or to cover work that is important but not continuous. A reasonable middle path is an agency build with a contractual handover, so your team inherits something documented rather than starting cold.

### What is the difference between an AI agency and an AI consultancy?

In practice the labels are used interchangeably and neither is protected. What matters is which end of the work a firm actually does. Some produce strategy and roadmaps and never touch a system. Some build and leave. The question to ask is whether the same firm is accountable for the strategy, the build, and the quality of the system six months later.

### How long does it take to get an AI system live?

A scoped first system commonly reaches a production pilot in six to twelve weeks, assuming access to content and systems is granted in the first fortnight. Access and data quality are usually the slow part, not the model work. Larger programmes are better run as a series of production releases than as one long build.

### Do we need our own model, or is fine-tuning necessary?

Usually neither. Most business problems are retrieval problems: the system needs to find the right passage of your content and answer from it with a citation. Retrieval is cheaper, updates the moment your content does, and produces answers a human can verify. Fine-tuning earns its place for format, style and narrow classification work. Be cautious when a supplier proposes it by default.

### What are the warning signs when choosing an AI agency?

A guaranteed outcome. A quote produced before anyone has looked at your content. No evaluation plan. No answer on permissions. A demonstration that only works on questions the supplier chose. Reluctance to name the people who will do the work. A proposal that never mentions what happens when the system is wrong. Any claim to be the best or the leading agency without evidence to support it.

### Does an AI agency need to be local to us?

Rarely. Almost all of this work is done remotely, and most UK firms serve clients across the country and abroad. Locality matters when you want people in the room for workshops, when procurement prefers a supplier in the same jurisdiction, or when data residency requirements make the supplier's location relevant. Ask where processing happens rather than where the office is.

### How do we check an AI agency is a real, accountable business?

Look the company up at Companies House. Check whether it is registered with the Information Commissioner's Office if it will handle personal data. Ask for evidence of insurance through procurement. Ask who signs the contract. These take minutes and rule out a surprising number of suppliers.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed 05 September 2026.

Keep reading

## Related

### AI Consulting & Development

What we actually do, and the method behind it, in the same plain terms.

Read this →

### AI agency in Cardiff

Where we are, who we work with in Wales, and what local actually buys you.

Read this →

### Working with London clients

How a firm with no London office serves London businesses, honestly described.

Read this →

## Put these questions to us

Send the one-page problem description described above. We will answer all nine questions in writing, and tell you if we think another kind of firm fits better.

Start a conversation → 
 See the live systems
