Free tool
LLM Routing & Cache Savings Estimator
Routing and caching get quoted as one savings number, which makes it impossible to tell which one is carrying it. They are separated here, because they carry very different risk.
Today, per month
$54,792
Everything on the current model, nothing cached.
After both levers
$21,761
60% lower.
From routing
$22,602
Carries quality risk. Every task class needs measuring before it moves.
From caching
$10,430
Carries no quality risk at all. This is the one to do first.
Annualised saving
$396,072
At today's traffic. Agent workloads rarely shrink, so treat this as a floor.
Runs entirely in your browser. Nothing you type here is transmitted, logged, or stored anywhere. There is no server-side copy because there is no transmission. Close the tab and it is gone.
How this works, and what it does not cover
Routing moves a share of calls to a cheaper model, and the saving is the price difference across those calls. It can cost you quality, so every task class needs measuring before it moves.
Caching applies a discount to the share of input tokens that repeat between calls — system prompt, tool definitions, conversation history. It cannot cost you quality at all, which is why it is almost always the thing to do first.
Cache discounts are provider-dependent and change. The default here is a common figure, not a promise; check your own pricing page.
What this does not model: batching, provider commitments, the engineering time to implement either lever, or the quality regression testing that routing honestly requires.
The routing decision table: which task classes move safely to a smaller model, which do not, and how the quality check was run.
If you would rather not do this yourself
LLM Cost & Routing Sprint
Three weeks: find where the money goes, fix the largest lines, and leave the controls that stop it coming back.
3 weeks · $15-20K