Free tool

Agent SLO & Error Budget Calculator

Reliability targets are cheap to declare and expensive to ignore. This turns four objectives into error budgets, and error budgets into the thresholds that decide when someone gets paged.

Window
days

Thirty days is the usual choice. A quarter hides too much.

Task completion rate
%

Tasks that reach a correct terminal state without human rescue.

%
Hallucination escape rate
%

Responses reaching a user with an error the system should have caught. Lower is better.

%
End-to-end latency
% within SLA
%
Tool-call budget

Not a reliability objective so much as a cost one, but it belongs in the same review.

Error budget over 30 days

60,000 tasks in the window. 3 of 3 objectives are already over budget.

Completion — budget consumed

200%

1,800 failed tasks allowed, 3,600 spent.

Hallucination escapes — budget consumed

280%

300 escapes allowed, 840 spent.

Latency — budget consumed

250%

600 slow tasks allowed, 1,500 spent.

Tool calls over budget

180,000

3.0 extra calls per task across the window.

Alerting thresholds these imply

  • Page someone when the budget for any objective is burning fast enough to exhaust the window in under a day — on the completion objective, that is roughly 75 failed tasks per hour, sustained for a full day.
  • Open a ticket at 50% of a budget consumed with more than half the window remaining.
  • Freeze risky releases at 100%. An error budget with no consequence attached is a dashboard, not a policy.

Runs entirely in your browser. Nothing you type here is transmitted, logged, or stored anywhere. There is no server-side copy because there is no transmission. Close the tab and it is gone.

How this works, and what it does not cover

An error budget is the volume of failure an objective permits over a window. A 97% completion objective across 60,000 tasks permits 1,800 failures — the budget is the number, and consuming it is not a crisis until it runs out.

Three of the four objectives are ordinary reliability engineering. Hallucination escape rate is the one teams argue about, because measuring it means deciding what counts as an error that should have been caught. That argument is usually where the product improves.

The tool-call budget is a cost objective rather than a reliability one, but it belongs in the same review: a system that drifts from 12 calls per task to 15 has changed in a way that shows up on the invoice before it shows up in the dashboards.

What this does not model: correlation between objectives, seasonality, or the fact that the same underlying defect often spends two budgets at once.

The SLO starter pack: the four objective definitions, the error budget policy template, and the alerting thresholds behind the 71% to 96% programme.

No cookies, no tracking pixels, never shared or sold. How the data is handled.

If you would rather not do this yourself

Reliability Foundation — SLOs and evals in CI

Three weeks: reliability targets your team agrees on, error budgets tied to them, and evaluations that run in CI rather than in someone's head.

3 weeks · $12-18K
All free tools