Production AI Systems

Systems that change a number, not pilots that prove a concept

We build AI into a named business process and redesign the process around it. Agentic patterns where the task genuinely requires planning and tool use. Deterministic automation where it does not. Human review by default.

Our position on agents

We will talk you out of an agent if you do not need one

17% of organisations have deployed AI agents. Over 40% of agentic projects are forecast to be cancelled by the end of 2027 on cost, unclear value or inadequate controls. Of the thousands of vendors claiming agentic capability, Gartner assesses roughly 130 as genuinely agentic.

Gartner, 2025–2026

We build agentic systems, and we are good at it. We also think most of what is sold as agentic should be a deterministic workflow with one model call in it, which is cheaper, faster and auditable. The baseline usually settles the argument with data.

Use an agent when

The path genuinely varies per case, the system must choose among tools, and a wrong step is recoverable and reviewable.

Do not when

The process is stable, the steps are known, the output is regulated, or nobody can articulate what “correct” looks like well enough to grade it.

Processes we have built into, or would take on

Claims and case handling

Intake, triage, evidence gathering and a drafted decision with the reasoning attached. Human sign-off retained.

Client reporting

Recurring reports assembled from systems of record, in house voice, with every figure traceable to its source.

Document production

Regulated correspondence and structured documents where the template is known and the content is not.

Back-office exceptions

The exception queue: mismatches, missing data, things that fell out of the happy path.

Underwriting support

Assembling the pack, flagging the anomalies, and never making the decision.

Knowledge and internal copilots

Useful, and honestly the hardest to attach a number to. We will say so before you fund it.

Every build ships with its own evidence

Evaluation suite

100+ example golden set, binary pass or fail, judge calibrated against human labels.

Acceptance thresholds

Agreed before build, not negotiated after. Below threshold means it does not ship.

Escalation design

What the system refuses to do alone, and who picks it up when it does.

Runbook

So your team can operate it, whether or not you keep us on to do it.
FAQs

Questions worth answering

When should a workflow use an AI agent, and when should it not?

Use an agent when the path genuinely varies per case, the system must choose among tools, and a wrong step is recoverable and reviewable. Do not use one when the process is stable, the steps are known, the output is regulated, or nobody can articulate what "correct" looks like well enough to grade it.

Are most "agentic" products actually agentic?

No. 17% of organisations have deployed AI agents, over 40% of agentic projects are forecast to be cancelled by the end of 2027 on cost, unclear value or inadequate controls, and of the thousands of vendors claiming agentic capability Gartner assesses roughly 130 as genuinely agentic. Much of what is sold as agentic should be a deterministic workflow with one model call in it, which is cheaper, faster and auditable.

What ships alongside a production AI system?

An evaluation suite with a 100+ example golden set graded pass or fail and a judge calibrated against human labels; acceptance thresholds agreed before the build; an escalation design defining what the system refuses to do alone; and a runbook so your own team can operate it.

Have a process in mind?

Bring us the one with the queue. We will baseline it before we quote to build anything.