Services · AI

AI agents that act, and the four things that make that safe

An assistant answers a question. An agent goes and does the thing, which is a much larger promise and a much larger liability. We build agentic AI for teams that need work completed rather than described, and we are candid about the tasks where a plain script is the better answer.

Tool-using agentsHuman in the loopEvery action logged
Quick answer

An AI agent is a language model given tools, memory and a loop, so it can carry out a multi-step task rather than describe one. Green Arrow Consultancy works as an AI automation agency: we design the tools an agent may call, scope them to least privilege, keep irreversible steps behind human confirmation, and log every action so a wrong one can be found and undone.

Definitions

What an AI agent actually is

Strip the marketing away and an agent is four things. A model, which supplies the judgement. A set of tools, which are ordinary functions with typed inputs and real side effects: create the ticket, look up the order, write the row, send the message. Memory, which is at minimum the transcript of what has happened so far in this run and often something carried between runs. And a loop, which lets the model call a tool, read what came back, and decide again.

The loop is the part that matters. In a normal language model call you get one shot: the model writes an answer and the exchange is over. In an agent, the model sees the consequence of its last decision before it makes the next one. That is what allows recovery. The record was not found, so try the other identifier. The API returned a validation error, so fix the field and retry. No script written in advance handles the second and third order cases the way a competent operator does, and that is the capability you are buying.

It is also what makes agents expensive and occasionally alarming. The same loop that recovers from a failure can pursue a misunderstanding for fifteen steps at full price. Every tool result is appended to the transcript, so the cost of step nine is higher than the cost of step two, and nobody decided that in advance: the model did, at runtime.

Assistant, workflow, agent

Three things get called automation and they behave nothing alike. A workflow is a fixed graph that you authored; it does exactly what you drew, forever. An assistant is a single call that returns text and changes nothing. An agent is a workflow whose next node is chosen by a model each time round. The engineering job is to give that chooser a small enough set of options that its worst decision is still one you can live with.

Most of what gets sold as an agent is a workflow with a language model in one of the boxes. That is a perfectly good thing to build, and it is often the right thing to build. It just should not be priced or governed as though it were an agent, and an agent should never be governed as though it were a workflow.

Judgement call

When an agent beats a script, and when it does not

We use this table in the first meeting. It has ended more agent projects than it has started, which is the point of having it.

Shape of the taskAgent?What to build instead, or alongside
The same steps, in the same order, every timeNoA scheduled job or a rule in the system you already own. Cheaper, deterministic, and it cannot improvise.
Moving structured data between two systems on a scheduleNoAn integration. An agent adds latency, cost and a new failure mode in exchange for nothing.
One question, one answer, drawn from documentsNoA retrieval assistant with citations. See AI consulting for how we build those.
High volume, tight latency, low value per runRarelyA classifier plus rules. Reserve the agent for the small fraction the classifier flags as unusual.
Unstructured input that has to be read before the route is knownYesAgent, with a narrow tool set and a confidence threshold that hands off to a person.
A task whose step count depends on what is found along the wayYesAgent. This is the case agents exist for, and the case scripts handle badly.
A long tail of requests that are similar but never identicalYesAgent, plus a routing layer so the routine ninety per cent never reaches it.
Anything irreversible: payments, deletions, messages to customersOnly with a humanAgent prepares, person commits. The commit action is not exposed to the model at all.

An honest agency turns down agent work. The failure mode of this market is an agent built where a two hundred line integration would have been better in every measurable way.

Safety

The four conditions for a safe agent

These are structural properties of the tool layer, not instructions in a prompt. A prompt is a request that a model may decline to follow. A permission boundary is not.

Engineering

Designing the loop and stopping runaway cost

The orchestration layer, not the model, owns the loop. It decides how many times round the agent may go, how much it may spend doing so, what counts as progress, and what happens when progress stops. Writing please use no more than ten steps in a system prompt is a hope. A counter in code is a rule.

The stopping conditions worth having

A hard step ceiling, sized from the evaluation runs rather than guessed. A token and cost budget per run, enforced before the call is made rather than reported afterwards. A wall clock limit, because an agent quietly retrying a timing-out API is a cost leak with no output. A repetition detector, since the commonest runaway is the same tool called with the same arguments four times in a row. And a no-progress rule: if the state of the world has not changed in three steps, stop and escalate with the transcript.

Context growth is the cost curve nobody models

Every tool result joins the transcript, so a long run pays for its own history repeatedly. Left alone, an agent that reads six large documents will spend most of its budget re-reading them. The fixes are unglamorous: summarise tool output at the boundary instead of pasting it whole, keep large payloads in a scratch store and pass references, prune completed sub-tasks out of the working transcript, and cache the stable parts of the prompt. These are the difference between an agent that costs pennies a run and one the finance team eventually switches off.

Not every step needs the biggest model

Planning and ambiguous judgement justify a frontier model such as Claude. Extracting a reference number from a returned payload does not. Routing steps by difficulty is the single largest cost reduction available in most agent designs, and because the application layer is kept separate from the model, changing that routing is configuration plus a re-run of the evaluation set.

Failure should be loud and clean

Agents fail closed. When a tool errors twice, when confidence drops below threshold, when the budget is spent, or when the task turns out to be outside scope, the run stops, rolls back what it can, and hands a human the full transcript with a plain statement of what was and was not completed. The outcome to design against is not the visible failure. It is the silent partial success, where three of five steps happened and nobody knows which three.

Evaluation

Evaluating agents is genuinely hard

Evaluating a question-answering system is tractable. One input, one output, a reference answer, a score. Evaluating an agent is a different problem, because the thing being judged is a trajectory: a sequence of decisions through a space that branches at every step. Two trajectories can both be correct and share almost no steps. One can look sensible the whole way and end with the wrong record updated.

So we score the end state first. For every test case, the expected condition of the target systems after a correct run is written down in advance, and the check queries those systems directly. The agent's own account of what it did is evidence, not proof; models are perfectly capable of reporting success on a call that returned an error.

Step-level scoring sits on top of that, because end state alone hides bad habits. We check whether the right tool was chosen, whether the arguments were sane, whether the agent recovered from an injected failure, whether it asked for confirmation where it was meant to, and whether it stopped instead of grinding on. An agent that reaches the right answer after nineteen steps and two accidental writes has not passed.

You need a fixture environment

None of this is possible against production. Agent evaluation needs a seeded copy of the target systems that is reset between cases, so the same test runs the same way tomorrow. Building that copy is often the largest single line in an agent project, and it is the line clients most often ask to cut. It is also the only thing that will let you upgrade a model without holding your breath.

Adversarial cases are not optional

An agent reads data that other people wrote: ticket bodies, email threads, supplier PDFs, web pages. Any of those can carry an instruction aimed at the model rather than the reader. Indirect prompt injection is the top entry in the OWASP Top 10 for LLM Applications for good reason, and an agent with tools turns it from a content problem into an actions problem. We test it deliberately, alongside data exfiltration attempts and tool argument abuse. That work is described in more detail under AI security.

Logging is the deliverable that saves you

Every run writes a structured record: correlation identifier, trigger and triggering identity, each tool call with arguments and results, model and prompt versions, token and cost totals, stop reason, and any human approval with the approver's name and timestamp. That log is how an incident becomes a ten minute investigation. It is also most of the evidence base for the oversight and record-keeping expectations in the EU AI Act and ISO/IEC 42001, which we cover under AI governance.

Method

How an agent engagement runs

Six phases. The first one saves clients the most money, and the fourth one is the one that gets argued about.

  1. 01

    Qualify the task, and be willing to say no

    We start with the shape of the work, not the technology. Does the path vary at runtime, or only the data? If the steps are fixed we recommend an integration and tell you what it should cost. Roughly half of the agent enquiries that reach us end here, which we regard as the process working.

  2. 02

    Design the tools before anything else

    Tool contracts come first: what each function does, its typed inputs, its narrowest useful scope, its credentials, whether it is reversible, and whether it needs human confirmation. The set of tools is the real security boundary, so it is designed with whoever owns those systems in the room.

  3. 03

    Engineer the loop

    Orchestration, planning strategy, stopping conditions, step and cost ceilings, context pruning, retry policy, escalation behaviour and the shape of the human approval step. This is where an agent stops being a demonstration.

  4. 04

    Build the evaluation harness

    A seeded fixture environment, cases drawn from real work, end-state assertions, step-level scoring, injected failures and adversarial prompts. The harness is written before autonomy is widened, not after the first incident.

  5. 05

    Pilot under supervision

    The agent runs in shadow mode or with confirmation on every action while its output is compared against the human baseline. Autonomy widens one action at a time, on evidence, and only for actions where the cost of a wrong call is known.

  6. 06

    Operate, monitor and review

    Dashboards for cost per completed task, failure and escalation rates, and human override frequency. Regression runs on every model or prompt change. A quarterly review that asks which confirmations can now be relaxed and which should be reinstated.

Honesty

What agents cannot do yet

Every one of these is a live research area and some will move within the year. None of them should be assumed away in a system going live next quarter.

Long-horizon reliability
Per-step accuracy compounds. An agent that gets each step right nineteen times in twenty is not reliable across twenty dependent steps, and no prompt fixes that arithmetic. Long tasks are decomposed into short supervised runs with checkpoints between them, or they are not built.
Judgement nobody wrote down
An agent can apply a policy. It cannot infer the unwritten rule that this particular customer is handled differently, or that the finance director wants the January figures presented a certain way. Where the knowledge lives only in people's heads, the first deliverable is writing it down, and that is a consulting job before it is an engineering one.
Noticing that the task is wrong
Agents are poor at stepping back. Given an instruction resting on a false premise, the usual behaviour is a confident, well-executed answer to the wrong question. Humans catch this because they carry context the agent was never given.
Resisting instructions hidden in data
Current models cannot reliably distinguish content they are reading from instructions they are following. Mitigation is architectural, through scoped tools, confirmation gates and output filtering, not something to be solved with a firmer system prompt.
Working reliably through a screen
Computer-use agents that drive a browser or a desktop are improving quickly and are still brittle against layout changes, sessions and consent dialogs. Where an API exists, we use the API. Where none exists, we say so, and price the risk instead of hiding it.
Being cheap while being open-ended
You may have open-ended exploration or predictable cost, and most designs are chosen by which of the two the business actually needs. Pretending both are available at once is how agent pilots quietly get switched off in month four.
In production

Where we run this ourselves

We would rather you tried the work than read about it. The demonstrations run on sample datasets, so nothing you type touches a real customer record.

Agent flows inside Neuro Search

A request is planned into steps, executed against the systems we have connected, and returned with every source cited so a person can check the work rather than trust it. The planning and the citation are the same feature: an agent whose steps you cannot inspect is an agent you cannot approve.

Document remediation at library scale

A production system that repairs PowerPoint and PDF libraries for accessibility, working file by file with a review stage before anything is written back. A clear example of prepare and commit as separate actions.

Operations work that agents suit

Reconciliation, exception handling, case triage and report assembly are where the step count genuinely varies. We keep a plain account of which operations tasks are worth automating and which are not.

Questions

Frequently asked questions

More on models, security and ways of working in the full FAQ, and the vocabulary is defined in the glossary.

What is the difference between an AI agent and a chatbot?

A chatbot produces text. An agent produces changes. The chatbot takes your question, retrieves something relevant and writes an answer; nothing in the world is different afterwards. An agent is given a set of tools, which are functions that touch real systems, and a loop that lets it call one, read the result and decide what to do next. When it finishes, a ticket exists, a record has moved, a file has been written or an email is sitting in a drafts folder. That difference is the whole engineering problem, because a wrong sentence is embarrassing and a wrong action is expensive.

When is an AI agent the wrong tool?

Whenever the steps are known in advance and never vary. If the work is always the same five actions in the same order, a script, a scheduled job or a rule in the system you already own will do it faster, cheaper and with no chance of improvisation. Agents earn their cost when the path is decided by what the agent finds along the way: unstructured input, several possible routes, a step count that is not known until the work starts. We turn down more agent requests than we accept, usually in favour of an integration.

How do you stop an agent doing something irreversible?

By making irreversible actions structurally impossible without a person. Destructive and outbound capabilities are either not exposed as tools at all, or are exposed as a two-stage pair: the agent may prepare and the human may commit. An agent can draft the customer email, stage the refund, or assemble the deletion list; a named person presses send. This is a property of the tool layer rather than an instruction in the prompt, because a prompt is a request and a permission boundary is a rule.

What does it cost to run an AI agent?

More than people expect, because the unit is not one model call. A single agent run is a sequence of calls whose length the model decides at runtime, and each call carries the whole growing transcript of tool results with it, so cost per step rises as the run goes on. We budget per run rather than per month, cap steps and tokens in the orchestration code, route the cheap steps to smaller models, and instrument cost per completed task so the finance question has an actual number attached to it.

How long does an agent project take?

A first agent in production usually takes eight to fourteen weeks, and the slow part is almost never the model. It is getting API access to the target systems, agreeing which actions a machine may take without a human, and building the fixture environment the agent can be tested against. Where the tools already exist and the permissions question is already settled, it is faster. Where neither is true, that work happens first whether or not anyone planned for it.

How do you evaluate an agent, when there is no single correct answer?

You score the end state, not the prose. Each test case defines what the target systems should look like after a correct run, and the check queries those systems rather than trusting the agent's own summary. Step-level scoring sits on top of that. Runs happen against a seeded copy of the target system that is reset between cases, because you cannot evaluate an agent by letting it loose on production.

Can an agent work inside the systems we already have?

That is the only version worth building. Agents that live in their own console get used for a fortnight. We put the work where the work already happens, which in practice means Salesforce, SharePoint, WordPress, Shopify, a ticketing system, a finance system or an internal application with an API. If a system has no API, that constraint shapes the design honestly rather than being solved with screen scraping we would have to apologise for later.

What can AI agents not do yet?

Long-horizon planning is the clearest limit, because per-step accuracy compounds and a twenty step task is far worse than twice a ten step one. Agents also cannot supply judgement nobody wrote down, cannot be relied on to notice that the premise of a task was wrong, and cannot resist an instruction hidden in the data they read unless the tool layer stops them.

Who is accountable when an agent gets something wrong?

You are, which is why the design has to make that survivable. Every run carries a correlation identifier, a full record of tool calls with arguments and results, the model and prompt versions in force, and the identity of whoever or whatever triggered it. That record is what lets someone reconstruct a bad outcome in minutes instead of arguing about it. It is also what an auditor, an insurer or an EU AI Act conformity assessment will ask to see.

Written and reviewed by the Green Arrow Consultancy team, led by Darren Tyler, founder and chief executive.

Green Arrow Consultancy Ltd, Cardiff, Wales. Company number 12491770. ICO registration ZA822868. Member of the International Association of Privacy Professionals. Last reviewed .

Tell us the task, and we will tell you if it needs an agent

Describe the work you want taken off someone's desk. We will say whether an agent is the right instrument, what the tool set would look like, and where a person still has to press the button.