Every agency in this market says the same six things. This guide is the set of questions that actually separate them, written by a firm that has to answer them too.
There is no single best AI agency in the UK, and any firm that claims the title should be asked to substantiate it. The better question is which agency fits your problem. Judge on four things: a system running in production you can look at, an evaluation set they will hand over, a clear answer on data handling, and a named person who does the work.
People search for the best AI agency in the UK
because they want a shortcut through a market with no comparable outputs. It is a reasonable instinct and there is no shortcut. Two agencies can both be excellent and be wrong for each other's clients, because the work ranges from data platform engineering to conversation design to regulated document processing, and almost nobody is genuinely strong across all of it.
There is also a rule worth knowing as a buyer. Under the UK advertising code, a claim that consumers would read as objective has to be backed by documentary evidence, and a superlative like best or number one is treated as a comparison against every competitor, which means it has to be verifiable. So when a supplier's homepage says they are the leading AI agency in the country
, the correct response is to ask what evidence they hold. Often there is none, and that tells you something about how the rest of the proposal was written.
What you can compare is evidence. Whether a system is actually running. Whether anyone measures its quality. What happens when it is wrong. Who is accountable. Those are answerable questions, and the answers vary enormously between suppliers who look identical on a website.
Ask all nine, of every supplier, in the same order. The pattern of what they answer easily and what they deflect is more informative than any single answer.
Not a slide, not a recorded video, a system a real organisation uses. Ask who uses it, how long it has been live, and what broke in the first month. An agency that cannot answer the third question has probably never operated anything.
There should be an immediate answer: a citation the user can check, a refusal path when retrieval finds nothing, a way for a support agent to escalate, and a log that lets someone reproduce it. If the answer is 'we tune the prompt', the system has no safety design.
Ask for an evaluation set: real questions, agreed correct answers, and a score that runs on every change. Ask whether you keep it at the end. Without one, quality is an opinion and you will have no way to tell when a vendor model update degrades your system.
An assistant that reads everything will eventually tell someone something they should not have seen. Ask how the permission model of your source systems is enforced inside retrieval, not around it. This is the most common gap in AI work built by teams without enterprise integration experience.
Ask which provider, which tier, whether training is contractually excluded, where processing happens, what is logged and for how long, and what appears in the record of processing. Ask for it in writing before the pilot, not before the contract.
Providers deprecate models on their own schedule. Ask whether the application layer is separate from the model, and what the swap actually costs. If the answer is a rewrite, you are buying a liability with a hidden expiry date.
Ask for the names of the people who will build it, not the names on the pitch. Ask what happens if that person leaves. In small firms you get continuity but less depth; in large ones the reverse. Both are fine, but know which you are buying.
Ask what you receive: source, infrastructure definitions, the evaluation set, the runbooks, the model card. An agency that will only rent you the system should say so plainly so you can price the lock-in.
The most useful question in the list. A supplier who has never talked a client out of an AI project either has not had the conversation or did not want to. A good answer names the conditions under which the project fails and offers to check them first.
| What the proposal says | What to check | What good looks like |
|---|---|---|
| "We use the latest models" | Which one, on which tier, and what happens when it is retired | A named model, a named tier, and an architecture that survives a swap |
| "The AI is trained on your data" | Whether they mean fine-tuning or retrieval, because the two are not alike | Retrieval for knowledge, with fine-tuning proposed only for format or style |
| "Enterprise-grade security" | The permission model inside retrieval, and where processing happens | A written data flow and a record of processing before the pilot starts |
| "We will iterate based on feedback" | Whether there is an evaluation set, and who keeps it | Real questions, agreed answers, a score that runs on every release |
| "Fully managed service" | What handover looks like if you leave | A documented exit: source, infrastructure, evaluation set, runbooks |
| A fixed price with no discovery | What it is based on, since nobody has read your content | A fixed-price discovery first, then a quote grounded in what was found |
This is a general guide to reading proposals. It is not a comment on any named firm.
Nobody can give you a credible figure for an AI system before looking at your content, and the ones who do are pricing a template. What can be described honestly is the shape of the cost and the things that move it.
Most engagements have three parts. A fixed-price discovery of two to three weeks that ends with a recommendation and a real quote. A build, priced per system once the discovery has established what the content and access actually look like. Then a monthly retainer for running it: evaluation, monitoring, model updates and change requests. If a proposal has no third part, ask who is responsible for quality in month seven.
Content that is out of date, duplicated or locked in formats nobody can parse. Source systems with no API. A permission model that has to be reconstructed before retrieval can respect it. Regulatory review. Multiple languages. A requirement for on-premise or in-region processing. Integration into a system whose owner has not agreed to the project.
A narrow first use case with a clear measure. Content that is already reasonably organised. A named internal owner who can get access granted in days rather than months. Willingness to ship something narrow and useful before something broad and impressive.
The single largest cost driver is not the technology. It is how long it takes to get access to the content, and that is usually determined on your side rather than the supplier's.
None of these mean a supplier is dishonest. They mean there is a question to ask before you sign.
A short paid comparison costs less than one wrong hire and produces evidence rather than proposals.
One page: the task, who does it today, and the measure that would say it worked.
Identical brief, identical questions, identical deadline. Different briefs make proposals incomparable.
A small fixed fee buys you their thinking on your real content, and their willingness to say no.
The supplier who arrives with an evaluation set has done this before.
Ask the supplier for it. The reaction to the request is itself informative.
Record the reason. It is what you will re-read when the project gets difficult.
This guide is published by Green Arrow Consultancy, which is one of the firms a reader might be comparing. So here is our position without adjectives.
We are a small UK firm, founded in 2012 and based in Cardiff, that came to AI from web development, privacy engineering and accessibility rather than from data science. That shapes what we are useful for. We are a reasonable choice when a system has to reach production inside an organisation that has compliance obligations, when retrieval and permissions matter more than novel modelling, and when the same firm handling the AI also needs to handle the website, the consent layer and the accessibility work around it.
We are the wrong choice if you need a large team mobilised quickly, bespoke model research, or a supplier who will put dozens of people on site. Those are real requirements and other firms serve them better than we would.
You can check the claims in this paragraph rather than take them: seven of our production systems are generalised into live demonstrations you can use without signing up, the company is registered at Companies House under number 12491770, and the about page links the registrations directly. Apply the nine questions above to us as well.
There is no single best AI agency in the UK, and any firm claiming the title should be asked for its evidence, because UK advertising rules require objective superlatives to be substantiated. The useful question is which agency is right for your problem. A firm that is excellent at large-scale data platform work may be the wrong choice for a customer assistant grounded in a product catalogue, and the reverse is equally true. Judge on production evidence, evaluation practice, data handling and who actually does the work.
Discovery is usually a fixed fee for two to three weeks of work. A first production system is typically quoted after discovery, because the honest range before anyone has looked at your content is too wide to be useful. Ongoing operation is normally a monthly retainer covering evaluation, model updates, monitoring and change requests. Beware any quote given before a supplier has seen your data, and any quote with no line for running the thing after launch.
Build in-house if AI is going to be core to your product and you can hire and keep the people. Use an agency to get the first systems into production, to bring a discipline you do not yet have, or to cover work that is important but not continuous. A reasonable middle path is an agency build with a contractual handover, so your team inherits something documented rather than starting cold.
In practice the labels are used interchangeably and neither is protected. What matters is which end of the work a firm actually does. Some produce strategy and roadmaps and never touch a system. Some build and leave. The question to ask is whether the same firm is accountable for the strategy, the build, and the quality of the system six months later.
A scoped first system commonly reaches a production pilot in six to twelve weeks, assuming access to content and systems is granted in the first fortnight. Access and data quality are usually the slow part, not the model work. Larger programmes are better run as a series of production releases than as one long build.
Usually neither. Most business problems are retrieval problems: the system needs to find the right passage of your content and answer from it with a citation. Retrieval is cheaper, updates the moment your content does, and produces answers a human can verify. Fine-tuning earns its place for format, style and narrow classification work. Be cautious when a supplier proposes it by default.
A guaranteed outcome. A quote produced before anyone has looked at your content. No evaluation plan. No answer on permissions. A demonstration that only works on questions the supplier chose. Reluctance to name the people who will do the work. A proposal that never mentions what happens when the system is wrong. Any claim to be the best or the leading agency without evidence to support it.
Rarely. Almost all of this work is done remotely, and most UK firms serve clients across the country and abroad. Locality matters when you want people in the room for workshops, when procurement prefers a supplier in the same jurisdiction, or when data residency requirements make the supplier's location relevant. Ask where processing happens rather than where the office is.
Look the company up at Companies House. Check whether it is registered with the Information Commissioner's Office if it will handle personal data. Ask for evidence of insurance through procurement. Ask who signs the contract. These take minutes and rule out a surprising number of suppliers.
Send the one-page problem description described above. We will answer all nine questions in writing, and tell you if we think another kind of firm fits better.