AI & Applied Machine Learning
AI with a scoreboard, a fallback and a human in the loop.
We are deliberately unromantic about this. AI earns its place when there is a repetitive, expensive, judgement-light task; enough historical examples to evaluate against; and a tolerable cost of being wrong. When those conditions hold, the returns are large and fast. When they don’t, a rules engine is cheaper, more accurate and easier to defend, and that is what we’ll recommend, even though it’s a smaller invoice.
- Feasibility sprints that end in “don’t build it”
- ~1/3Feasibility sprints that end in “don’t build it”
- Every generated answer traceable to a source
- CitedEvery generated answer traceable to a source
- In the loop until accuracy earns autonomy
- HumanIn the loop until accuracy earns autonomy
Sounds like
You might recognise one of these.
Four people spend their week retyping data from PDFs.
We tried a chatbot and it confidently made things up.
Our forecast is a spreadsheet someone maintains by hand.
The board wants an AI strategy and we need it to be real.
What this includes
The work, specifically.
Not every engagement needs all of it. This is the range we cover and what each part is actually for.
Document intelligence
Extraction from invoices, bills of lading, claims, EOBs, purchase orders and scanned contracts, with confidence thresholds that route uncertain cases to a person instead of guessing.
Retrieval-augmented assistants
Question answering grounded strictly in your own documents, with citations, refusal behavior when the source is silent, and access control enforced at retrieval time rather than in the prompt.
Classification, routing and prioritization
Triage of tickets, claims, leads and exceptions, deployed behind a queue so the model’s suggestion is reviewed until it has earned autonomy.
Forecasting and optimization
Demand, capacity, maintenance windows and route or schedule optimization, evaluated honestly against the naive baseline, because a surprising number of models lose to it.
Evaluation and guardrails
Every model ships with a labelled evaluation set, tracked accuracy and cost per task, prompt-injection defenses, output validation and a documented fallback for the day the provider has an outage.
What you get
Deliverables, not documents.
- A labelled evaluation set owned by your domain experts
- Measured accuracy, latency and cost per task against a baseline
- Human-in-the-loop review interface with correction capture
- Model- and vendor-agnostic abstraction with a defined fallback
- Prompt-injection and data-leakage threat review
- Monitoring for drift and cost, alerting to an owner
Shapes
How this usually runs.
AI feasibility sprint
2–3 weeksFixed fee. We take one candidate use case, build an evaluation set and prove or disprove it on your real data. Roughly a third of these end in a recommendation not to build, and that’s a good outcome.
Production pilot
6–12 weeksOne workflow, one team, in production behind a review queue, with the accuracy and savings measured rather than asserted.
Scale-out
OngoingAdditional use cases on shared evaluation, observability and governance infrastructure.
Tooling
What we build it with.
No tool here was picked because it was new. Where we do reach for something novel, it is in one place, for a stated reason, and it is written down.
- Models
- Retrieval
- Classical ML
- Operations
Questions
AI & ML, honestly.
Not unless you decide it should be. We deploy on enterprise terms with training explicitly disabled, or self-host open-weight models inside your own boundary when policy or classification requires it.
Constrain the task, ground answers in retrieved sources, validate outputs against a schema, require citations, and let the system say it doesn’t know. Then measure how often it’s wrong on a fixed evaluation set, because unmeasured accuracy is just a feeling.
We instrument cost per task from day one and design for it: smaller models where they suffice, caching, batching. Unit economics is a design constraint in these projects, not a surprise at the end of the month.
Further reading
What we think about this, at length.
- Data7 min read
In applied machine learning, the model is the tiebreaker
We built the full pipeline from raw network capture to a live detector, measured where the time went, and the answer was not the modeling.
- AI3 min read
AI with a scoreboard: which projects pay for themselves
Three conditions predict almost every applied-AI success we’ve shipped. Projects missing any one of them tend to fail in the same way.
Next step
Tell us what’s breaking.
Forty-five minutes, no charge, no deck. We’ll tell you what we’d do, what it would likely cost, and whether you should be building this at all.