LLM cost audit
We cut LLM bills for teams running AI in production.
You shipped AI features and they work. Now the model bill climbs every month and nobody can say what's driving it.
Where the money goes
Most teams have one bill and a guess.
You have a single charge from your model provider every month and a hope the numbers work out. That stops working the moment usage scales — and the invoice becomes the only signal you've got.
How it runs
Instrument, map, cut, prove.
Read-only access throughout. Your data stays in your environment — nothing leaves your infrastructure.
Capture the calls
Read-only instrumentation of your model and tool calls, inside your own environment.
See the spend
Cost broken down by feature, by model, and by call pattern — including the loops quietly burning tokens.
Make the changes
Route cheaper models where quality allows, kill wasteful retries, cache repeats. Each change checked against an eval so quality holds.
Hand back the number
Before and after, in writing. Was this, now that — with the eval showing output quality unchanged.
What you get
A spend map you've never actually seen — and a smaller bill.
- →A breakdown of where every pound of model spend goes, by feature and call pattern.
- →The routing and caching changes implemented, not just recommended.
- →A single before/after number, with the eval that proves quality held.
If we can't find safe savings, you don't pay for the implementation — only the audit.
Proof
Per-agent model routing on a production agent stack.
Each agent scoped to the cheapest model that held quality on its actual task. Spend down by over a third, output unchanged.
Anonymised. Full write-up on request.
Who it's for
Teams with LLM features live in production and a model bill big enough that a 20–40% cut is worth two weeks of someone's time.
Get started
Find out what your AI features actually cost.
A short call to see if there's spend worth cutting. No pitch deck.