Most AI programmes die between the pilot and the people who were meant to use it. We build the unglamorous half, the retrieval, the permissions, the evaluation and the governance, because that is the half that decides whether anything survives contact with a real business.
Green Arrow Consultancy is a UK AI agency that designs, builds and runs production AI systems. We work on retrieval assistants, AI agents, document AI and enterprise search for consumer brands, manufacturers and professional firms, with privacy engineering, evaluation and governance built in rather than added later. Founded in 2012, based in Cardiff, working across the UK, USA, EU and Asia Pacific.
There are two failure modes in this market. The first is the strategy firm that produces a maturity assessment, a roadmap and an opportunity matrix, and never touches a system. The second is the build shop that wires a language model to a chat box, demonstrates it on cherry-picked questions, and leaves before anyone asks what happens when the answer is wrong.
We think an AI agency is accountable for the whole line: deciding whether the problem deserves AI at all, getting access to the content that makes an answer possible, building the retrieval and permission layers that make it correct and safe, integrating it where the work already happens, and then staying on to watch quality over time. That last part is not optional. Language models are updated by their vendors, your content changes weekly, and a system that was accurate in March will quietly drift by September unless somebody is running the evaluation set.
It also means telling clients when the answer is no. A large share of the requests that reach us are better solved by fixing a search index, rewriting a policy document, or removing a form field. We would rather say so in week one than bill for a model that papers over it.
Where we came from matters here. Green Arrow Consultancy did not start as an AI company. We started in 2012 building and managing websites, then spent a decade running privacy, consent, analytics and accessibility programmes for global consumer brands across more than two hundred websites. That history is why our AI work looks the way it does: we already knew what a data protection impact assessment costs, what an accessibility audit finds, and what happens when a marketing team ships something the legal team has not seen.
Six shapes of work cover almost everything we are asked for. Each one exists in production somewhere, and each one is generalised into a demonstration you can open in a new tab.
Concierge-grade assistants grounded in your products, policies and tone of voice. Guard-railed so they stay on topic, cited so a human can verify, and escalated to a person the moment the conversation needs one.
One question across HR systems, procurement, policy libraries, file stores and custom internal tools. Permission-aware, so people only ever see answers drawn from documents they are entitled to read.
Systems that plan a multi-step task and execute it: raise the ticket, draft the reply, assemble the report, reconcile the list. Every action is logged, reversible and scoped to what the agent is allowed to touch.
Thousands of contracts, reports, manuals and decks turned into a corpus that answers questions with the page reference attached. Includes ingestion, chunking, embedding and the quality work that makes retrieval land.
A photograph of a part becomes the exact SKU. A floor plan becomes a bill of materials. A slide deck becomes an accessible, screen-reader-ready document with real alt text.
Evaluation sets built from your real questions, adversarial testing against prompt injection and data exfiltration, refusal behaviour verified rather than assumed, and regression runs on every release.
Five phases. The first and the last are the ones that decide the outcome.
We start with the decision or the task, not the technology. Who does this work today, what does good look like, and what measure will tell us it worked. If the honest answer is that a better search index or a rewritten policy solves it, that is the recommendation you get.
A narrow prototype against your real content and a set of the questions your people actually ask. This is where most of the uncomfortable discoveries happen, and it is much cheaper to have them here.
Retrieval design, chunking strategy, permission mapping, guardrails, escalation paths, integration into the tools people already have open. The model is one component of several, and it is rarely the component that decides whether the project succeeds.
Data flow and record of processing, model card, human oversight design, retention rules, an escalation route for bad answers, and the documentation that a regulator, an auditor or a nervous general counsel will eventually ask to see.
Monitoring, cost control, regression evaluation on every change, drift checks as vendor models update, and a quarterly review against the measure we agreed in phase one. This is the phase most suppliers skip.
After a decade of running other people's digital estates, the pattern is consistent enough to plan around.
| Stage | What people expect to be hard | What is actually hard |
|---|---|---|
| Discovery | Choosing the right model | Getting access to the content, and finding out how much of it is out of date |
| Prototype | Prompt engineering | Agreeing what a correct answer even looks like, in writing |
| Build | The AI layer | Permissions, so the system cannot answer from documents the asker may not read |
| Launch | Infrastructure | Change management, and the first three wrong answers a senior person sees |
| Operation | Cost | Silent quality drift, because nobody owns the evaluation set |
Observed across our own deployments and the programmes we have been brought in to rescue.
We are deliberately model-agnostic. Most production systems we run are built on Anthropic's Claude models, chosen for instruction following, long-context behaviour and their willingness to refuse rather than improvise. Cheaper and faster models are routed the simpler traffic. Where data residency, unit cost at scale or offline operation makes a hosted model impossible, we deploy open-weight models on infrastructure you control.
The important architectural decision is not which model. It is that the application layer, the retrieval layer and the evaluation layer are separate from the model, so that swapping the model is a configuration change and a re-run of the evaluation set rather than a rebuild. Vendors deprecate models on their own schedule. Any system that cannot survive that is a liability you have not been told about yet.
Fine-tuning is the answer far less often than it is proposed. Retrieval-augmented generation, where the system searches your content and answers from what it finds, is cheaper, updates the instant your content does, and produces citations a human can check. We reserve fine-tuning for genuine format and style problems, and we say so rather than selling the more expensive option by default.
An assistant that can read everything will eventually tell someone something they should not have seen. We map the permission model of the source systems into retrieval itself, so the corpus a question is answered from is already filtered to that person's entitlements. This is the single most common gap we find in AI systems built by teams without an enterprise integration background.
Every system we hand over comes with an evaluation set: real questions, agreed correct answers, and a harness that scores the system against them. It is what lets you upgrade a model without holding your breath, and it is what turns 'the AI seems worse lately' into a number you can act on.
The list is the same whether the engagement is six weeks or a year.
Green Arrow Consulting has been a strategic partner in the development and maintenance of the Energizer International Digital platform, working with our regional offices in Asia, Middle East and Africa and Europe.
We would rather you tried the work than read about it. Seven live systems, no sign-up, running on sample datasets so nothing you type touches a real customer record.
A browser extension that answers across HR, procurement, policies and custom internal systems, with citations and permission awareness. Built from an enterprise deployment, generalised for the demo.
A customer assistant grounded in a product catalogue, policies and brand voice, switchable across seven industries so you can see it answer in your own sector.
A customer maps their building zone by zone and the system returns a specified, orderable parts list with a branded PDF at the end. This one began as an industrial power advisor.
If your question is not here, the full FAQ covers privacy, security, pricing and ways of working.
An AI agency takes a business problem, decides whether AI is the right instrument for it, then designs, builds, evaluates and operates the system that solves it. In practice that means data and access work, retrieval design, prompt and model engineering, evaluation harnesses, guardrails, integration into the systems people already use, and the ongoing operations that keep quality from drifting. A consultancy that only produces strategy decks is not an AI agency; a development shop that ships an unevaluated chatbot is not one either.
A scoped first system typically takes six to twelve weeks from kick-off to a production pilot, assuming we can get access to the content and systems in the first fortnight. Discovery and access are usually the slow part, not the model work. Larger programmes run as a sequence of production releases rather than one long build, so value lands early and the direction can change on evidence.
We are model-agnostic and choose per workload. Most of our production systems run on Anthropic's Claude models because of their instruction following and long-context behaviour, with smaller and cheaper models routed the simpler tasks. Where data residency, cost at scale or offline operation demands it, we deploy open-weight models on infrastructure you control. The application layer is built so the model underneath can be swapped without a rewrite.
Not on the configurations we deploy. We use enterprise API tiers with training opt-out, keep retrieval corpora inside your tenancy or ours under contract, and document the data flow in a record of processing before anything goes live. Privacy engineering is the practice this firm grew out of, so this is a starting condition rather than an afterthought.
Three layers. Grounding, so the model answers from retrieved passages of your own content rather than memory. Citation, so every claim is traceable to the document it came from and a human can check it in one click. Refusal, so the system says it does not know when retrieval returns nothing relevant, which is a behaviour we test for explicitly. On top of that sits an evaluation set built from your real questions that runs on every change.
Discovery engagements are fixed price and typically run two to three weeks. Build work is quoted per system after discovery, because the honest range for 'an AI assistant' is too wide to be useful before anyone has looked at the content. Ongoing operation is a monthly retainer covering evaluation, model updates, monitoring and change requests. We will tell you when a problem does not need AI, which is more often than the market admits.
Yes, and a good share of our work is exactly that. We have spent over a decade as the third party that sits between in-house teams and their agencies, and we are used to being the specialist layer rather than the owner of the whole relationship. Where an internal team wants to take the system over, we build it to be handed over and document it accordingly.
Consumer products and retail, industrial and manufacturing, hospitality, healthcare, financial services, education and legal. The published live demonstrations on this site switch between those seven sectors because each one came out of a real deployment before it was generalised.
Send us the task you want removed from someone's week. We will tell you whether AI is the right instrument, roughly what it costs, and what has to be true for it to work.