Support & Run

The monthly report that says whether your AI still works

You built it, you bought it, or a platform vendor deployed it. A support and run contract puts each named production system under continuous evaluation and sends you one report a month: whether output quality has held, what it is costing, and what broke. The panel below is that report, in the format you receive it.

The monthly report

What the report tells you before you have to ask

A mock of our own reporting format, filled with sample data, for one AI system over one month. Pass rate is the share of sampled cases that met the agreed definition of correct. Cost per case is inference spend divided by the work done. The dip is a supplier changing model version without saying so, which is the kind of event this contract exists to catch.

claims-triage-prodHealthyLast 30 days · updated 4 min ago
Pass rate94.2% threshold 90% · +1.4pt
Cases handled18,402 +7% on prior period
Cost / case£0.031 budget £0.05 · caching on
Open incidents1 SEV-3 · drift on intake form v4
Evaluation pass rate, 30 days pass rate threshold
1 Aug18 Aug · model version change detected, rolled back in 41 min30 Aug

Interface mock. Figures are illustrative sample data, not a client system.

Somebody has to own whether it still works

That is what the contract is. No vague retainer: a severity model, response times, and a metric we report against whether it flatters us or not.

SEV-1 · 1 hrSEV-2 · 4 hrsSEV-3 · next working day

Continuous evaluation

Every release and a sampled share of live traffic scored against agreed acceptance thresholds, with the judge itself calibrated against human labels.

Drift, including judge drift

Input distribution drift, silent model version changes by your supplier, and movement in the evaluator’s own behaviour. The third is the one most people never watch.

Cost control

Token spend per unit of work, budget alerts, caching and batching, and a recommendation each quarter on whether a cheaper model would hold quality.

Incident response

Defined severities, a named responder, rollback procedure, and a written post-incident note that goes in your audit trail.

Quarterly improvement

One release a quarter aimed at the metric in your baseline, with the change in that metric reported rather than the work delivered.

Service levels

Scoped by the number of production systems under contract and the severity cover they need.

Watch

One production system. Monitoring, evaluation, monthly report, SEV-2 response.

Operate

Up to three systems. Adds SEV-1 response, cost optimisation and a quarterly improvement release.

Estate

Four or more systems, or a regulated environment needing named oversight. Audit and certification support is scoped separately with Pixelette Certified.

How it starts

It does not have to be something we built

We will take on systems we did not build, once a baseline tells us what we are inheriting.

FAQs

Questions worth answering

How is an AI support and run contract priced?

Against the number of production systems under contract and the severity cover they need, not against headcount or hours. Watch covers one production system with monitoring, evaluation, a monthly report and SEV-2 response. Operate covers up to three systems and adds SEV-1 response, cost optimisation and a quarterly improvement release. Estate covers four or more systems, or a regulated environment needing named oversight. Every contract is scoped and quoted after a conversation.

What are the response times?

SEV-1 within 1 hour, SEV-2 within 4 hours, SEV-3 by the next working day. Each incident gets a named responder, a rollback procedure and a written post-incident note for your audit trail.

Will Pixelette support an AI system it did not build?

Yes, once a baseline establishes what is being inherited. Running what somebody else wrote is the clearest proof that this is a capability rather than a warranty on our own work.

What is judge drift?

Movement in the behaviour of the model doing the grading, as distinct from drift in the input data or a silent model version change by a supplier. Almost nobody watches it, and it quietly invalidates your quality measurements when it happens.

Already have something in production?

The baseline works just as well on a system that exists as on one that does not. We measure what it is doing now and tell you what it would cost to keep it honest.