Automate · Agentic AI

Autonomy is a control decision, not a slogan

Single- and multi-agent systems that plan, call tools, coordinate steps and operate within defined controls. Human oversight is the default where decisions are material; more autonomous operation is introduced where the workflow, the risk and the evidence justify it.

What an agentic build actually consists of

Very little of it is prompt writing. The engineering lives in what the system may reach, what it does when a step fails, and how the run is reconstructed afterwards.

Planning and task decomposition

Turning an instruction into an ordered set of steps the system can actually carry out, with the plan visible and inspectable rather than implied by whatever the model felt like doing.

Tool use and system access

The functions, queries and APIs an agent may call, what arguments it may pass, and the explicit list of things it may not touch. The boundary is the design.

Multi-agent coordination

Where a task genuinely splits into research, drafting, checking and execution, agents with distinct responsibilities and a defined protocol between them, rather than one prompt pretending to be a team.

State, memory and recovery

What the system remembers within a task and across tasks, and what it does when a step fails halfway through. Recovery design is most of the engineering.

Human approval points

The moments where the run stops and a person decides. Placed where the consequence is material, not where they are least inconvenient.

Traceability

Every step, tool call, input and decision recorded, so an outcome can be explained afterwards to a customer, an auditor or a court.

How autonomy is granted

Four steps, and you have to earn each one

A system moves up only on evidence from its own evaluation results. Nothing here is granted because a demonstration went well.

01

Suggest

The system proposes; a person does the work. Nothing changes without a human action. The right place to start with almost anything.
02

Draft for approval

The system produces the output and a person approves or edits it before it leaves the building. Most production value sits here.
03

Act within bounds

The system executes inside a narrow, reversible envelope, escalating anything outside it. Requires evaluation evidence before it is granted.
04

Act and report

The system runs and a person reviews after the fact. Justified only by measured performance over time on a task where a mistake is recoverable.

The same ladder runs in reverse. If measured quality falls below the agreed threshold, the system drops a step and a person is back in the loop until it recovers. That is a designed behaviour, not an incident response.

Where this fits

An agent is a solution to variability, not to work

Agentic architecture earns its cost when the path genuinely varies from case to case and the system has to choose among tools to get through it. When the steps are known and stable, the same outcome is available from a deterministic workflow with a model call inside it. That is cheaper to run, faster, and very much easier to explain to whoever asks why it did what it did.

We build both, and we would rather argue for the simpler one before you have paid for the complicated one.

How we build production AI

What has to exist first

An agent needs somewhere to act and something to act on. In practice that means the tools it calls have to exist as reliable interfaces, the data it reads has to be reachable with the right permissions, and there has to be an agreed definition of a correct outcome precise enough to grade against.

Where those are missing, the honest sequence is integration first, evaluation second, agent third. Reversing it produces a demonstration rather than a system.

FAQs

Questions worth answering

What is a multi-agent system, in practical terms?

A design in which distinct components each hold a defined responsibility, for example gathering information, drafting, checking and executing, and communicate through a defined protocol. It is worth the extra complexity only where the task genuinely splits along those lines. Where it does not, a single agent, or a deterministic workflow with one model call in it, is cheaper, faster and easier to audit.

How do you decide how much autonomy an agent should have?

Autonomy is treated as an engineered control decision rather than a selling point. Human oversight is the default wherever a decision is material. More autonomous operation is introduced only where the workflow, the risk level and measured evaluation evidence justify it, in defined steps: suggest, draft for approval, act within bounds, then act and report.

How do you stop an agent doing something it should not?

By constraining what it can reach rather than by asking it nicely. The tools, queries and APIs available to an agent are an explicit allow-list, arguments are validated, irreversible actions are gated behind human approval, and every step and tool call is recorded so the run can be reconstructed afterwards.

Thinking about agents?

Bring us the process and we will tell you, with the reasons, whether it needs an agent, a deterministic workflow, or a fortnight of integration work first.