Services · AI

AI chatbot development, where the content is the actual job

The chat window is a week of work. Everything that decides whether the assistant is trusted, the retrieval, the citations, the refusals, the escalation and the state of your documents, is the rest of the project. We build that half.

Grounded in your contentEvery answer citedEscalates to a person
Quick answer

Green Arrow Consultancy builds custom AI chatbots that answer from your own content and cite where each answer came from. We are a UK AI chatbot development company working on public website assistants, permission-aware internal assistants and support-agent copilots. Every build ships with guardrails, refusal behaviour, an escalation route to a person, an accessibility pass and an evaluation set you keep.

The honest version

What a custom AI chatbot actually is

Custom means two things, and neither is the colour of the launcher button. The assistant answers from your content rather than from what the model absorbed in training, and its behaviour is designed for your business: what it discusses, what it declines, when it fetches a person.

The technology is not exotic. A question arrives, a search runs across your documents, the relevant passages go in front of a language model, and it answers from them with references attached. Anyone competent can wire up retrieval-augmented generation in a fortnight.

Which is why the wiring is not the job. The content behind the assistant is the job. Grounded in nothing, an assistant has only the model's memory, so it invents plausible answers and cannot show where they came from. Grounded in a corpus where three documents disagree about your returns window, it answers differently depending on what the search surfaces.

Citations are a product feature, not a technicality

An answer that can be checked earns the provisional trust a new system needs. An uncited assistant that is wrong starts an argument about whether the machine is any good. A cited one that is wrong produces a link to the document that is wrong: a content ticket, not a crisis.

Refusal has to be built on purpose

Models are obliging by default and will attempt an out-of-scope question rather than decline. Saying that is not in the material I can see, here is how to reach the team who has it has to be designed, prompted and verified in the evaluation set. If you are comparing suppliers, choosing an AI agency in the UK lists the questions worth asking all of us.

Scope

Three things people mean by chatbot

These are not variations of one product. They differ in who reads them, what they may see and how they fail.

Type of assistantWho it is right forWhat it connects toThe risk that defines it
A customer-facing assistant on a public websiteBrands with a product range, a policy library and a support queue full of the same twenty questions.Product data, help centre articles, delivery, returns and warranty policies, order status where an API exists.It is public. Everything it says is a statement by your company, and anybody can spend an afternoon talking it out of its instructions.
An internal assistant over company systemsOrganisations where the answer exists but nobody can find it: HR, operations, compliance, field teams.SharePoint and intranet pages, policy libraries, process manuals, resolved tickets, HR systems, internal applications.Permissions. An assistant that reads everything will eventually tell someone something they were not entitled to see, which is a reportable event.
A support-agent copilot inside the helpdeskService teams with high volume, long onboarding and a knowledge base agents have stopped opening.The ticketing system, canned responses, resolved tickets, product and policy content.Agent confidence. Suggest two wrong answers in one morning and the panel is ignored for the rest of the year. Adoption is the measure, not accuracy.

Many organisations end up with two of the three. Scope and launch them separately: they fail in different directions.

Deliverables

What gets built

The same list whether the assistant answers twenty questions or twenty thousand.

Reality check

Where chatbot projects fail

We have rescued enough of these to see the pattern. Almost none of the failures are model failures.

The content is stale, or it contradicts itself

The commonest one by a distance. The help centre says fourteen days, the terms say thirty, a PDF from two years ago says something else. A human reading all three picks the right one from context. Retrieval returns whichever passage scored highest. Every organisation believes its content is broadly correct until something reads all of it at once.

Nobody owns the content

The assistant produces a stream of questions it could not answer. That list is valuable and worth nothing if no named person acts on it. We ask for that name during discovery.

There is no escalation path

An assistant deployed to reduce contact volume, with no route to a person behind it, does not reduce it. It moves the conversation somewhere angrier, usually a public review. If there is no human queue, decide that consciously before launch.

There is no evaluation, so quality is opinion

Without an evaluation set the only evidence is anecdote. Someone senior tries three questions, one is poor, and the project acquires a reputation quiet improvement never shifts. With one, quality is a number that moved or did not.

The scope is too broad on day one

The instinct is to answer everything, because a narrow assistant seems underwhelming. Broad scope means a corpus too large to have been checked, and a good chance the first thing a director asks is the thing it handles worst. Narrow first, widen on evidence.

The classic: a public launch with no guardrails

A model on a public chat window with no scope limits, no refusal behaviour and no injection testing will be made to say something embarrassing, usually within days, by someone who then screenshots it. It will also discuss your competitors and offer opinions your legal team never approved. That is not a model defect. It is an assistant shipped without the half of the work that constrains it, which is why guardrails and adversarial testing are launch blockers here. If one is live already and it worries you, tell us what it did.

Method

How a build runs

Six stages. The second is the one clients want to skip, and the one that decides the rest.

  1. 01

    Decide which assistant this is, and what it must not do

    Public, internal or agent copilot. Then the unglamorous part: the questions it must answer, the topics it must decline, and what a correct answer looks like, in writing.

  2. 02

    Audit the content before writing any code

    An inventory of sources, a contradiction report, a list of documents with no owner, and a judgement about which corpus is defensible enough for a first release.

  3. 03

    Build retrieval and prove it on real questions

    Ingestion, chunking, hybrid search and reranking, tuned against questions your people have actually asked rather than ones invented for a demonstration. Retrieval is scored separately from answer quality: when an answer is wrong you need to know which half failed.

  4. 04

    Wrap it in guardrails, permissions and escalation

    Scope enforcement, refusal wording, permission filtering, prompt injection defences, output checks and the helpdesk handover. This decides whether the assistant is safe to make public.

  5. 05

    Design the interface and integrate it

    The chat surface, in your brand, on the page or inside the agent console. Accessibility is engineered here, not audited afterwards: retrofitting live regions into a streaming widget is a rebuild.

  6. 06

    Pilot supervised, launch, then keep scoring it

    A narrow audience first, conversations reviewed daily, the evaluation set growing from what turns up. Then a staged launch, monitoring, regression runs, and unanswered questions routed to the content owner.

Accessibility

The accessibility problem nobody mentions

Chat interfaces fail assistive technology users routinely, and it stays invisible because almost nobody tests it. We audit every interface we build against WCAG 2.2 AA, which matters under the European Accessibility Act and EN 301 549. Detail in accessible AI interfaces and accessibility.

Streaming answers a screen reader never hears
A chat assistant streams tokens into a container already on the page. Without a correctly configured live region, and a decision about announcing as it arrives or once complete, a screen reader user gets silence. Announcing every token produces an unusable stutter.
Focus that goes nowhere when the panel opens
The launcher is pressed, a dialog appears, and focus stays on the button behind it. The panel needs a dialog role, an accessible name, focus moved in on open and returned on close, and Escape that closes it.
Keyboard traps and unreachable controls
Suggested-reply chips that are divs rather than buttons, a send control reachable only by mouse, citation links that are not links. WCAG 2.2 criterion 2.1.2 is what fails.
Small targets, poor contrast and no reflow
Chat widgets are drawn small on purpose, which collides with WCAG 2.2 target size, contrast on placeholder text, and reflow at 400 per cent zoom.
Sessions that expire without warning
A conversation that clears itself after inactivity, with no warning and no way to extend, fails the timing criteria.
Language that assumes sight and speed
Instructions such as click the icon below, errors carried only by colour, progress shown only by animation.
Security and privacy

Security, privacy and what reaches the model

A public assistant is an input field connected to a language model, published on the internet with your brand on it. It brings one problem application security has no pattern for.

Prompt injection, in both directions

Current models cannot reliably separate content they are reading from instructions they are following. That arrives directly, from a visitor trying to override the system prompt, and indirectly, from an instruction hidden in a retrieved document or supplier PDF. The OWASP Top 10 for LLM Applications ranks prompt injection first, and the defences are architectural: strict scoping, output filtering, no privileged tools behind a public assistant, retrieval treated as untrusted input, and adversarial cases in the evaluation set. That testing sits under AI security; an assistant allowed to act rather than only answer is covered under AI agents.

What data actually reaches the model

The question worth putting to any supplier is precise: which text leaves your infrastructure, where does it go, who holds it, for how long. The model provider receives the question, the retrieved passages and the conversation so far. It does not receive your document store. We use enterprise API tiers with training opt-out and record the data flow before launch.

Transcripts are personal data more often than teams assume

People type order numbers, email addresses and complaints into chat windows regardless of what the interface asked for. Retention needs a decision and a period, the privacy notice needs to cover the chat, and a data protection impact assessment is appropriate where the processing warrants one under UK GDPR. That runs through privacy consulting and AI consulting.

Say that it is a machine

Assistants that imply a human is typing generate complaints, and the EU AI Act expects people to be told when they are interacting with an AI system. People forgive a machine that admits what it is far more readily than one that pretended.

Proof

Assistants you can open right now

We would rather you tried the work than read a claim about it. Sample data, no sign-up, and worth pushing off topic.

AI Concierge

A customer assistant grounded in a product catalogue, policies and brand voice, switchable across seven industries. It ran in production before it was generalised.

Customer assistantOpen the assistant →

Product finding by conversation

A visitor describes a situation rather than a specification, and the assistant asks what a good salesperson would before recommending.

Guided sellingTry the finder →

Comparison and specification answers

Side-by-side comparison from structured product data, with the answer explaining the difference rather than listing attributes.

Darren’s practical mind makes him really effective at cutting through complexity to find the most suitable solution.

NicoloSenior Brand Manager, Energizer Holdings
Questions

Frequently asked questions

More on models, data handling and ways of working in the full FAQ, with the vocabulary defined in the glossary.

How long does it take to build?

A focused first assistant usually reaches a supervised pilot in around six to ten weeks, and the slow part is almost never the model. It is getting access to the content, discovering how much of it contradicts itself, and agreeing who owns the corrections. Broader rollouts run as a sequence of releases.

Will it make things up?

Any language model can produce a confident wrong answer, so the job is to make that rare, visible and quick to correct rather than to promise it away. We ground answers in retrieved passages of your own content, cite every claim so a reader can check it, and test refusal so the assistant says it does not know when retrieval finds nothing relevant. Be wary of anyone who says the number will be perfect.

Can it use our own documents?

That is the normal case and the reason to build a custom assistant at all. Policies, product data, manuals, help centre articles and past tickets become a retrieval corpus: ingested, split sensibly, indexed for vector and keyword search. Nothing is fine-tuned into the model, so a corrected document is live as soon as the index updates.

Can it hand over to a human?

Yes, and the handover is designed before the conversation flow is. It escalates on request, on repeated failure, on out-of-scope topics such as complaints or safety, and on evident frustration. What matters is what travels with it: the transcript, what was retrieved and what could not be found.

Is it GDPR compliant?

Compliance is a property of the deployment rather than of the software, so we build it in and document it: a lawful basis identified before launch, a data protection impact assessment where the processing warrants one, a record of processing, enterprise model tiers with training opt-out, deliberate transcript retention, and a clear notice that people are talking to an assistant. See privacy consulting.

Can you improve a chatbot we already have?

Often, and it is a common way for people to reach us. We sample real conversations, score them against what a correct answer would have been, and establish whether the fault is the model, the retrieval, the content underneath or a missing escalation route. Usually it is the content and the retrieval, which is good news: neither requires replacing the platform.

Do we need our own model?

Almost certainly not. Fine-tuning is proposed far more often than it is warranted, and it is the heavy answer to a retrieval problem. Retrieval keeps knowledge in documents you can edit, so corrections are immediate and citations are possible. We are model-agnostic, with most production assistants on hosted frontier models such as Claude, and open-weight models where data residency demands it.

Can it answer in more than one language?

Yes, and the constraint is the corpus rather than the model. An answer translated on the fly from English source material is no longer grounded in anything a local team approved, which matters where product claims or legal wording differ by market. We either ground each language in its own reviewed content, or restrict the assistant to the languages the corpus covers.

Do we have to fix our content before we start?

Not before, but during, and it is the part clients most often underestimate. The readiness stage tells you which documents contradict each other, which have no owner and which are out of date. We narrow the first release to a scope where the content is defensible. That audit is useful even if the assistant is never built.

What happens after launch?

Launch is when the useful data starts arriving. We review the conversations nobody could answer, which is the most valuable content backlog you will ever be handed, and send the gaps to whoever owns the source documents. Alongside that runs monitoring of escalation and refusal rates and regression scoring on every change. An assistant accurate in March will drift by September unless somebody runs the harness.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Tell us what people keep asking you

Send us the twenty questions your team answers every week and a link to where the answers are meant to live. We will tell you whether that content can carry an assistant yet.