An assistant answers a question. An agent goes and does the thing, which is a much larger promise and a much larger liability. We build agentic AI for teams that need work completed rather than described, and we are candid about the tasks where a plain script is the better answer.
An AI agent is a language model given tools, memory and a loop, so it can carry out a multi-step task rather than describe one. Green Arrow Consultancy works as an AI automation agency: we design the tools an agent may call, scope them to least privilege, keep irreversible steps behind human confirmation, and log every action so a wrong one can be found and undone.
Strip the marketing away and an agent is four things. A model, which supplies the judgement. A set of tools, which are ordinary functions with typed inputs and real side effects: create the ticket, look up the order, write the row, send the message. Memory, which is at minimum the transcript of what has happened so far in this run and often something carried between runs. And a loop, which lets the model call a tool, read what came back, and decide again.
The loop is the part that matters. In a normal language model call you get one shot: the model writes an answer and the exchange is over. In an agent, the model sees the consequence of its last decision before it makes the next one. That is what allows recovery. The record was not found, so try the other identifier. The API returned a validation error, so fix the field and retry. No script written in advance handles the second and third order cases the way a competent operator does, and that is the capability you are buying.
It is also what makes agents expensive and occasionally alarming. The same loop that recovers from a failure can pursue a misunderstanding for fifteen steps at full price. Every tool result is appended to the transcript, so the cost of step nine is higher than the cost of step two, and nobody decided that in advance: the model did, at runtime.
Three things get called automation and they behave nothing alike. A workflow is a fixed graph that you authored; it does exactly what you drew, forever. An assistant is a single call that returns text and changes nothing. An agent is a workflow whose next node is chosen by a model each time round. The engineering job is to give that chooser a small enough set of options that its worst decision is still one you can live with.
Most of what gets sold as an agent is a workflow with a language model in one of the boxes. That is a perfectly good thing to build, and it is often the right thing to build. It just should not be priced or governed as though it were an agent, and an agent should never be governed as though it were a workflow.
We use this table in the first meeting. It has ended more agent projects than it has started, which is the point of having it.
| Shape of the task | Agent? | What to build instead, or alongside |
|---|---|---|
| The same steps, in the same order, every time | No | A scheduled job or a rule in the system you already own. Cheaper, deterministic, and it cannot improvise. |
| Moving structured data between two systems on a schedule | No | An integration. An agent adds latency, cost and a new failure mode in exchange for nothing. |
| One question, one answer, drawn from documents | No | A retrieval assistant with citations. See AI consulting for how we build those. |
| High volume, tight latency, low value per run | Rarely | A classifier plus rules. Reserve the agent for the small fraction the classifier flags as unusual. |
| Unstructured input that has to be read before the route is known | Yes | Agent, with a narrow tool set and a confidence threshold that hands off to a person. |
| A task whose step count depends on what is found along the way | Yes | Agent. This is the case agents exist for, and the case scripts handle badly. |
| A long tail of requests that are similar but never identical | Yes | Agent, plus a routing layer so the routine ninety per cent never reaches it. |
| Anything irreversible: payments, deletions, messages to customers | Only with a human | Agent prepares, person commits. The commit action is not exposed to the model at all. |
An honest agency turns down agent work. The failure mode of this market is an agent built where a two hundred line integration would have been better in every measurable way.
These are structural properties of the tool layer, not instructions in a prompt. A prompt is a request that a model may decline to follow. A permission boundary is not.
The orchestration layer, not the model, owns the loop. It decides how many times round the agent may go, how much it may spend doing so, what counts as progress, and what happens when progress stops. Writing please use no more than ten steps in a system prompt is a hope. A counter in code is a rule.
A hard step ceiling, sized from the evaluation runs rather than guessed. A token and cost budget per run, enforced before the call is made rather than reported afterwards. A wall clock limit, because an agent quietly retrying a timing-out API is a cost leak with no output. A repetition detector, since the commonest runaway is the same tool called with the same arguments four times in a row. And a no-progress rule: if the state of the world has not changed in three steps, stop and escalate with the transcript.
Every tool result joins the transcript, so a long run pays for its own history repeatedly. Left alone, an agent that reads six large documents will spend most of its budget re-reading them. The fixes are unglamorous: summarise tool output at the boundary instead of pasting it whole, keep large payloads in a scratch store and pass references, prune completed sub-tasks out of the working transcript, and cache the stable parts of the prompt. These are the difference between an agent that costs pennies a run and one the finance team eventually switches off.
Planning and ambiguous judgement justify a frontier model such as Claude. Extracting a reference number from a returned payload does not. Routing steps by difficulty is the single largest cost reduction available in most agent designs, and because the application layer is kept separate from the model, changing that routing is configuration plus a re-run of the evaluation set.
Agents fail closed. When a tool errors twice, when confidence drops below threshold, when the budget is spent, or when the task turns out to be outside scope, the run stops, rolls back what it can, and hands a human the full transcript with a plain statement of what was and was not completed. The outcome to design against is not the visible failure. It is the silent partial success, where three of five steps happened and nobody knows which three.
Evaluating a question-answering system is tractable. One input, one output, a reference answer, a score. Evaluating an agent is a different problem, because the thing being judged is a trajectory: a sequence of decisions through a space that branches at every step. Two trajectories can both be correct and share almost no steps. One can look sensible the whole way and end with the wrong record updated.
So we score the end state first. For every test case, the expected condition of the target systems after a correct run is written down in advance, and the check queries those systems directly. The agent's own account of what it did is evidence, not proof; models are perfectly capable of reporting success on a call that returned an error.
Step-level scoring sits on top of that, because end state alone hides bad habits. We check whether the right tool was chosen, whether the arguments were sane, whether the agent recovered from an injected failure, whether it asked for confirmation where it was meant to, and whether it stopped instead of grinding on. An agent that reaches the right answer after nineteen steps and two accidental writes has not passed.
None of this is possible against production. Agent evaluation needs a seeded copy of the target systems that is reset between cases, so the same test runs the same way tomorrow. Building that copy is often the largest single line in an agent project, and it is the line clients most often ask to cut. It is also the only thing that will let you upgrade a model without holding your breath.
An agent reads data that other people wrote: ticket bodies, email threads, supplier PDFs, web pages. Any of those can carry an instruction aimed at the model rather than the reader. Indirect prompt injection is the top entry in the OWASP Top 10 for LLM Applications for good reason, and an agent with tools turns it from a content problem into an actions problem. We test it deliberately, alongside data exfiltration attempts and tool argument abuse. That work is described in more detail under AI security.
Every run writes a structured record: correlation identifier, trigger and triggering identity, each tool call with arguments and results, model and prompt versions, token and cost totals, stop reason, and any human approval with the approver's name and timestamp. That log is how an incident becomes a ten minute investigation. It is also most of the evidence base for the oversight and record-keeping expectations in the EU AI Act and ISO/IEC 42001, which we cover under AI governance.
Six phases. The first one saves clients the most money, and the fourth one is the one that gets argued about.
We start with the shape of the work, not the technology. Does the path vary at runtime, or only the data? If the steps are fixed we recommend an integration and tell you what it should cost. Roughly half of the agent enquiries that reach us end here, which we regard as the process working.
Tool contracts come first: what each function does, its typed inputs, its narrowest useful scope, its credentials, whether it is reversible, and whether it needs human confirmation. The set of tools is the real security boundary, so it is designed with whoever owns those systems in the room.
Orchestration, planning strategy, stopping conditions, step and cost ceilings, context pruning, retry policy, escalation behaviour and the shape of the human approval step. This is where an agent stops being a demonstration.
A seeded fixture environment, cases drawn from real work, end-state assertions, step-level scoring, injected failures and adversarial prompts. The harness is written before autonomy is widened, not after the first incident.
The agent runs in shadow mode or with confirmation on every action while its output is compared against the human baseline. Autonomy widens one action at a time, on evidence, and only for actions where the cost of a wrong call is known.
Dashboards for cost per completed task, failure and escalation rates, and human override frequency. Regression runs on every model or prompt change. A quarterly review that asks which confirmations can now be relaxed and which should be reinstated.
Every one of these is a live research area and some will move within the year. None of them should be assumed away in a system going live next quarter.
We would rather you tried the work than read about it. The demonstrations run on sample datasets, so nothing you type touches a real customer record.
A request is planned into steps, executed against the systems we have connected, and returned with every source cited so a person can check the work rather than trust it. The planning and the citation are the same feature: an agent whose steps you cannot inspect is an agent you cannot approve.
A production system that repairs PowerPoint and PDF libraries for accessibility, working file by file with a review stage before anything is written back. A clear example of prepare and commit as separate actions.
Reconciliation, exception handling, case triage and report assembly are where the step count genuinely varies. We keep a plain account of which operations tasks are worth automating and which are not.
More on models, security and ways of working in the full FAQ, and the vocabulary is defined in the glossary.
A chatbot produces text. An agent produces changes. The chatbot takes your question, retrieves something relevant and writes an answer; nothing in the world is different afterwards. An agent is given a set of tools, which are functions that touch real systems, and a loop that lets it call one, read the result and decide what to do next. When it finishes, a ticket exists, a record has moved, a file has been written or an email is sitting in a drafts folder. That difference is the whole engineering problem, because a wrong sentence is embarrassing and a wrong action is expensive.
Whenever the steps are known in advance and never vary. If the work is always the same five actions in the same order, a script, a scheduled job or a rule in the system you already own will do it faster, cheaper and with no chance of improvisation. Agents earn their cost when the path is decided by what the agent finds along the way: unstructured input, several possible routes, a step count that is not known until the work starts. We turn down more agent requests than we accept, usually in favour of an integration.
By making irreversible actions structurally impossible without a person. Destructive and outbound capabilities are either not exposed as tools at all, or are exposed as a two-stage pair: the agent may prepare and the human may commit. An agent can draft the customer email, stage the refund, or assemble the deletion list; a named person presses send. This is a property of the tool layer rather than an instruction in the prompt, because a prompt is a request and a permission boundary is a rule.
More than people expect, because the unit is not one model call. A single agent run is a sequence of calls whose length the model decides at runtime, and each call carries the whole growing transcript of tool results with it, so cost per step rises as the run goes on. We budget per run rather than per month, cap steps and tokens in the orchestration code, route the cheap steps to smaller models, and instrument cost per completed task so the finance question has an actual number attached to it.
A first agent in production usually takes eight to fourteen weeks, and the slow part is almost never the model. It is getting API access to the target systems, agreeing which actions a machine may take without a human, and building the fixture environment the agent can be tested against. Where the tools already exist and the permissions question is already settled, it is faster. Where neither is true, that work happens first whether or not anyone planned for it.
You score the end state, not the prose. Each test case defines what the target systems should look like after a correct run, and the check queries those systems rather than trusting the agent's own summary. Step-level scoring sits on top of that. Runs happen against a seeded copy of the target system that is reset between cases, because you cannot evaluate an agent by letting it loose on production.
That is the only version worth building. Agents that live in their own console get used for a fortnight. We put the work where the work already happens, which in practice means Salesforce, SharePoint, WordPress, Shopify, a ticketing system, a finance system or an internal application with an API. If a system has no API, that constraint shapes the design honestly rather than being solved with screen scraping we would have to apologise for later.
Long-horizon planning is the clearest limit, because per-step accuracy compounds and a twenty step task is far worse than twice a ten step one. Agents also cannot supply judgement nobody wrote down, cannot be relied on to notice that the premise of a task was wrong, and cannot resist an instruction hidden in the data they read unless the tool layer stops them.
You are, which is why the design has to make that survivable. Every run carries a correlation identifier, a full record of tool calls with arguments and results, the model and prompt versions in force, and the identity of whoever or whatever triggered it. That record is what lets someone reconstruct a bad outcome in minutes instead of arguing about it. It is also what an auditor, an insurer or an EU AI Act conformity assessment will ask to see.
Describe the work you want taken off someone's desk. We will say whether an agent is the right instrument, what the tool set would look like, and where a person still has to press the button.