The chat window is a week of work. Everything that decides whether the assistant is trusted, the retrieval, the citations, the refusals, the escalation and the state of your documents, is the rest of the project. We build that half.
Green Arrow Consultancy builds custom AI chatbots that answer from your own content and cite where each answer came from. We are a UK AI chatbot development company working on public website assistants, permission-aware internal assistants and support-agent copilots. Every build ships with guardrails, refusal behaviour, an escalation route to a person, an accessibility pass and an evaluation set you keep.
Custom means two things, and neither is the colour of the launcher button. The assistant answers from your content rather than from what the model absorbed in training, and its behaviour is designed for your business: what it discusses, what it declines, when it fetches a person.
The technology is not exotic. A question arrives, a search runs across your documents, the relevant passages go in front of a language model, and it answers from them with references attached. Anyone competent can wire up retrieval-augmented generation in a fortnight.
Which is why the wiring is not the job. The content behind the assistant is the job. Grounded in nothing, an assistant has only the model's memory, so it invents plausible answers and cannot show where they came from. Grounded in a corpus where three documents disagree about your returns window, it answers differently depending on what the search surfaces.
An answer that can be checked earns the provisional trust a new system needs. An uncited assistant that is wrong starts an argument about whether the machine is any good. A cited one that is wrong produces a link to the document that is wrong: a content ticket, not a crisis.
Models are obliging by default and will attempt an out-of-scope question rather than decline. Saying that is not in the material I can see, here is how to reach the team who has it has to be designed, prompted and verified in the evaluation set. If you are comparing suppliers, choosing an AI agency in the UK lists the questions worth asking all of us.
These are not variations of one product. They differ in who reads them, what they may see and how they fail.
| Type of assistant | Who it is right for | What it connects to | The risk that defines it |
|---|---|---|---|
| A customer-facing assistant on a public website | Brands with a product range, a policy library and a support queue full of the same twenty questions. | Product data, help centre articles, delivery, returns and warranty policies, order status where an API exists. | It is public. Everything it says is a statement by your company, and anybody can spend an afternoon talking it out of its instructions. |
| An internal assistant over company systems | Organisations where the answer exists but nobody can find it: HR, operations, compliance, field teams. | SharePoint and intranet pages, policy libraries, process manuals, resolved tickets, HR systems, internal applications. | Permissions. An assistant that reads everything will eventually tell someone something they were not entitled to see, which is a reportable event. |
| A support-agent copilot inside the helpdesk | Service teams with high volume, long onboarding and a knowledge base agents have stopped opening. | The ticketing system, canned responses, resolved tickets, product and policy content. | Agent confidence. Suggest two wrong answers in one morning and the panel is ignored for the rest of the year. Adoption is the measure, not accuracy. |
Many organisations end up with two of the three. Scope and launch them separately: they fail in different directions.
The same list whether the assistant answers twenty questions or twenty thousand.
We have rescued enough of these to see the pattern. Almost none of the failures are model failures.
The commonest one by a distance. The help centre says fourteen days, the terms say thirty, a PDF from two years ago says something else. A human reading all three picks the right one from context. Retrieval returns whichever passage scored highest. Every organisation believes its content is broadly correct until something reads all of it at once.
The assistant produces a stream of questions it could not answer. That list is valuable and worth nothing if no named person acts on it. We ask for that name during discovery.
An assistant deployed to reduce contact volume, with no route to a person behind it, does not reduce it. It moves the conversation somewhere angrier, usually a public review. If there is no human queue, decide that consciously before launch.
Without an evaluation set the only evidence is anecdote. Someone senior tries three questions, one is poor, and the project acquires a reputation quiet improvement never shifts. With one, quality is a number that moved or did not.
The instinct is to answer everything, because a narrow assistant seems underwhelming. Broad scope means a corpus too large to have been checked, and a good chance the first thing a director asks is the thing it handles worst. Narrow first, widen on evidence.
A model on a public chat window with no scope limits, no refusal behaviour and no injection testing will be made to say something embarrassing, usually within days, by someone who then screenshots it. It will also discuss your competitors and offer opinions your legal team never approved. That is not a model defect. It is an assistant shipped without the half of the work that constrains it, which is why guardrails and adversarial testing are launch blockers here. If one is live already and it worries you, tell us what it did.
Six stages. The second is the one clients want to skip, and the one that decides the rest.
Public, internal or agent copilot. Then the unglamorous part: the questions it must answer, the topics it must decline, and what a correct answer looks like, in writing.
An inventory of sources, a contradiction report, a list of documents with no owner, and a judgement about which corpus is defensible enough for a first release.
Ingestion, chunking, hybrid search and reranking, tuned against questions your people have actually asked rather than ones invented for a demonstration. Retrieval is scored separately from answer quality: when an answer is wrong you need to know which half failed.
Scope enforcement, refusal wording, permission filtering, prompt injection defences, output checks and the helpdesk handover. This decides whether the assistant is safe to make public.
The chat surface, in your brand, on the page or inside the agent console. Accessibility is engineered here, not audited afterwards: retrofitting live regions into a streaming widget is a rebuild.
A narrow audience first, conversations reviewed daily, the evaluation set growing from what turns up. Then a staged launch, monitoring, regression runs, and unanswered questions routed to the content owner.
Chat interfaces fail assistive technology users routinely, and it stays invisible because almost nobody tests it. We audit every interface we build against WCAG 2.2 AA, which matters under the European Accessibility Act and EN 301 549. Detail in accessible AI interfaces and accessibility.
A public assistant is an input field connected to a language model, published on the internet with your brand on it. It brings one problem application security has no pattern for.
Current models cannot reliably separate content they are reading from instructions they are following. That arrives directly, from a visitor trying to override the system prompt, and indirectly, from an instruction hidden in a retrieved document or supplier PDF. The OWASP Top 10 for LLM Applications ranks prompt injection first, and the defences are architectural: strict scoping, output filtering, no privileged tools behind a public assistant, retrieval treated as untrusted input, and adversarial cases in the evaluation set. That testing sits under AI security; an assistant allowed to act rather than only answer is covered under AI agents.
The question worth putting to any supplier is precise: which text leaves your infrastructure, where does it go, who holds it, for how long. The model provider receives the question, the retrieved passages and the conversation so far. It does not receive your document store. We use enterprise API tiers with training opt-out and record the data flow before launch.
People type order numbers, email addresses and complaints into chat windows regardless of what the interface asked for. Retention needs a decision and a period, the privacy notice needs to cover the chat, and a data protection impact assessment is appropriate where the processing warrants one under UK GDPR. That runs through privacy consulting and AI consulting.
Assistants that imply a human is typing generate complaints, and the EU AI Act expects people to be told when they are interacting with an AI system. People forgive a machine that admits what it is far more readily than one that pretended.
We would rather you tried the work than read a claim about it. Sample data, no sign-up, and worth pushing off topic.
A customer assistant grounded in a product catalogue, policies and brand voice, switchable across seven industries. It ran in production before it was generalised.
A visitor describes a situation rather than a specification, and the assistant asks what a good salesperson would before recommending.
Side-by-side comparison from structured product data, with the answer explaining the difference rather than listing attributes.
Darren’s practical mind makes him really effective at cutting through complexity to find the most suitable solution.
More on models, data handling and ways of working in the full FAQ, with the vocabulary defined in the glossary.
A focused first assistant usually reaches a supervised pilot in around six to ten weeks, and the slow part is almost never the model. It is getting access to the content, discovering how much of it contradicts itself, and agreeing who owns the corrections. Broader rollouts run as a sequence of releases.
Any language model can produce a confident wrong answer, so the job is to make that rare, visible and quick to correct rather than to promise it away. We ground answers in retrieved passages of your own content, cite every claim so a reader can check it, and test refusal so the assistant says it does not know when retrieval finds nothing relevant. Be wary of anyone who says the number will be perfect.
That is the normal case and the reason to build a custom assistant at all. Policies, product data, manuals, help centre articles and past tickets become a retrieval corpus: ingested, split sensibly, indexed for vector and keyword search. Nothing is fine-tuned into the model, so a corrected document is live as soon as the index updates.
Yes, and the handover is designed before the conversation flow is. It escalates on request, on repeated failure, on out-of-scope topics such as complaints or safety, and on evident frustration. What matters is what travels with it: the transcript, what was retrieved and what could not be found.
Compliance is a property of the deployment rather than of the software, so we build it in and document it: a lawful basis identified before launch, a data protection impact assessment where the processing warrants one, a record of processing, enterprise model tiers with training opt-out, deliberate transcript retention, and a clear notice that people are talking to an assistant. See privacy consulting.
Often, and it is a common way for people to reach us. We sample real conversations, score them against what a correct answer would have been, and establish whether the fault is the model, the retrieval, the content underneath or a missing escalation route. Usually it is the content and the retrieval, which is good news: neither requires replacing the platform.
Almost certainly not. Fine-tuning is proposed far more often than it is warranted, and it is the heavy answer to a retrieval problem. Retrieval keeps knowledge in documents you can edit, so corrections are immediate and citations are possible. We are model-agnostic, with most production assistants on hosted frontier models such as Claude, and open-weight models where data residency demands it.
Yes, and the constraint is the corpus rather than the model. An answer translated on the fly from English source material is no longer grounded in anything a local team approved, which matters where product claims or legal wording differ by market. We either ground each language in its own reviewed content, or restrict the assistant to the languages the corpus covers.
Not before, but during, and it is the part clients most often underestimate. The readiness stage tells you which documents contradict each other, which have no owner and which are out of date. We narrow the first release to a scope where the content is defensible. That audit is useful even if the assistant is never built.
Launch is when the useful data starts arriving. We review the conversations nobody could answer, which is the most valuable content backlog you will ever be handed, and send the gaps to whoever owns the source documents. Alongside that runs monitoring of escalation and refusal rates and regression scoring on every change. An assistant accurate in March will drift by September unless somebody runs the harness.
Send us the twenty questions your team answers every week and a link to where the answers are meant to live. We will tell you whether that content can carry an assistant yet.