Reading, listening and looking, where the use case supports it
Extraction, classification, sentiment, speech and image or video intelligence. That closing qualifier is doing real work: these systems are excellent at tasks with a consistent correct answer and unreliable at tasks without one.
What we build
Document extraction
Pulling structured fields out of invoices, forms, contracts, statements and correspondence, including the scanned, rotated and photographed versions that arrive in real post rooms.
Classification and routing
Deciding what a message, case or document is and where it should go. Usually the highest-value language task in a business, and the least glamorous.
Sentiment and theme analysis
What customers are actually saying across reviews, tickets and calls, aggregated into themes a team can act on rather than a single number nobody acts on.
Speech to text and voice interfaces
Transcription, diarisation, call summarisation and voice-driven interaction, with accent and audio-quality performance measured on your recordings rather than a vendor benchmark.
Image and video intelligence
Detection, classification, condition assessment and quality checking, where the visual task is defined tightly enough for a consistent answer to exist.
Multilingual handling
Detection, translation and language-specific behaviour, with the accuracy stated per language instead of averaged into a flattering headline.
Four rules that decide whether it survives contact with real inputs
Define correct first
Two people labelling the same hundred documents will disagree. Where they disagree, no system can be graded, so the definition is settled before the build.Confidence drives routing
High-confidence output flows through; low-confidence output goes to a person. That threshold is the main control you have, and it is yours to set.Measured on your material
Performance on a public benchmark tells you very little about performance on your forms, your accents and your photographs.Accuracy stated per class
An overall figure hides the category that matters. We report per class, including the rare ones, because that is where the cost of an error usually lives.None of this is exotic. It is the ordinary discipline of measuring a system against a labelled set built from real cases, and it is what separates a capability from a demonstration.
Some of these tasks should not be automated at all
Inferring emotion, intent, character or truthfulness about a person from their voice, face or writing is a category we treat with considerable caution: the evidence base is weak, the failure modes fall unevenly across groups, and the consequences of being wrong land on an individual.
Sentiment about a product in a review is a different proposition from a judgement about a person in an interview, and we will draw that line explicitly rather than quietly build whatever was asked for.
The reliable wins are unglamorous: extracting fields from documents that currently get typed in twice, classifying and routing inbound work that a person triages by hand, and summarising calls that otherwise generate a note nobody writes.
Each of those has a countable current cost, which means the value is provable rather than asserted.
Questions worth answering
What language, speech and vision work do you take on?
Document extraction, classification and routing, sentiment and theme analysis, speech to text with diarisation and summarisation, image and video intelligence, and multilingual handling. The qualifier that matters is whether the use case supports it: the task has to be defined tightly enough that a consistent correct answer exists.
How accurate will an extraction or classification system be?
That cannot honestly be answered before it is measured on your own material, because performance on public benchmarks says little about performance on your forms, accents and photographs. The engagement establishes accuracy per class against a labelled set built from real cases, and the confidence threshold that routes uncertain output to a person is set from those results.
What happens to the cases the system is unsure about?
They go to a person. Confidence-based routing is the primary control in this category: high-confidence output flows through, low-confidence output is queued for human review with the reason attached, and the volume in that queue is reported so the threshold can be tuned deliberately rather than discovered.
Have a pile of documents, calls or images?
Send us a description of what is in it and what someone currently does with it by hand. We will tell you whether the task is well enough defined to automate.