LLM cost audit

We cut LLM bills for teams running AI in production.

You shipped AI features and they work. Now the model bill climbs every month and nobody can say what's driving it.

Fixed scope · two weeks · read-only access

Monthly model spend
£48,200
after audit −£17,300 · −36%
Illustrative. Typical reduction 20–40%, no quality loss.

Where the money goes

Most teams have one bill and a guess.

You have a single charge from your model provider every month and a hope the numbers work out. That stops working the moment usage scales — and the invoice becomes the only signal you've got.

01
Top-tier models doing cheap work. The most expensive model runs tasks a smaller one handles fine.
02
Agent loops burning tokens. A task makes ten model calls where two would do, and nobody notices until the invoice.
03
No attribution. You can't tell which feature, model, or call pattern is driving the spend.

How it runs

Instrument, map, cut, prove.

Read-only access throughout. Your data stays in your environment — nothing leaves your infrastructure.

01 — INSTRUMENT

Capture the calls

Read-only instrumentation of your model and tool calls, inside your own environment.

02 — MAP

See the spend

Cost broken down by feature, by model, and by call pattern — including the loops quietly burning tokens.

03 — CUT

Make the changes

Route cheaper models where quality allows, kill wasteful retries, cache repeats. Each change checked against an eval so quality holds.

04 — PROVE

Hand back the number

Before and after, in writing. Was this, now that — with the eval showing output quality unchanged.

Two weeks. Week one: measure. Week two: cut and report.

What you get

A spend map you've never actually seen — and a smaller bill.

  • A breakdown of where every pound of model spend goes, by feature and call pattern.
  • The routing and caching changes implemented, not just recommended.
  • A single before/after number, with the eval that proves quality held.
No saving, no fee

If we can't find safe savings, you don't pay for the implementation — only the audit.

Proof

−36%

Per-agent model routing on a production agent stack.

Each agent scoped to the cheapest model that held quality on its actual task. Spend down by over a third, output unchanged.

Anonymised. Full write-up on request.

Who it's for

Teams with LLM features live in production and a model bill big enough that a 20–40% cut is worth two weeks of someone's time.

Get started

Find out what your AI features actually cost.

A short call to see if there's spend worth cutting. No pitch deck.