Reliability Foundation — SLOs and evals in CI
Every incident is argued over opinion instead of data.
Three weeks: reliability targets your team agrees on, error budgets tied to them, and evaluations that run in CI rather than in someone's head.
3 weeks · $12-18K · Signed off by CTO or VP Engineering
71% to 96%
SLO compliance across three client platforms, two quarters
99.9%
Uptime sustained across production containerized environments
900ms to 120ms
p99 query latency, through index and materialized-view redesign
Who this is for
- Your team ships agent features and cannot agree on whether a release made things better or worse.
- QA is vibes-based: someone tries a few prompts and calls it fine.
- You want to say something specific about reliability to a customer or a board.
And who it is not
- You want an observability tool installed. Most of the tooling is open source and straightforward to stand up — the engagement is the layer that makes it mean something: objectives, error budgets, and evals wired into CI.
- You have no production traffic yet. There is nothing to set an objective against.
What you get
- SLO definitions on task completion rate, mean tool calls, hallucination escape rate, and end-to-end latency.
- Error budgets tied to each objective, with the policy for what happens when one is spent.
- Agent trace instrumentation feeding dashboards your team will actually open.
- An evaluation suite running in CI, with regression diffs on every pull request.
- Alerting and runbooks for the failures that will actually page someone.
How it runs
Week 1
Define
What working means, agreed with the people who will be held to it. Objectives that cannot be measured get cut in this week, not defended for a quarter.
Week 2
Instrument
Traces, dashboards, and the evaluation suite wired into CI so regressions surface on the pull request rather than in production.
Week 3
Operate
Error budget policy, alerting thresholds, runbooks, and a dry run of an incident with your on-call.
This has shipped
03
SLO framework program
SLO compliance from 71% to 96% over two quarters.
07
Real-time trading system
Sub-10ms latency at peak. Slippage cut 14%, zero limit breaches on real capital.
08
Inference platform at scale
Cold starts 18 min to under 4. p99 from 900ms+ to 120ms. 99.9% uptime.
Read the code first
What it costs
$12-18K
Fixed for the agreed scope. Moves with the number of services instrumented.
Every engagement can start as a $2,500 paid pilot, credited in full against the full scope if you go ahead.
Every engagement and what it costsBefore you talk to anyone
Agent SLO & Error Budget Calculator
Turn reliability targets into error budgets, alert thresholds, and a number your team can argue about before the incident. Free, runs in your browser, and nothing you type leaves it.
Open the agent slo & error budget calculatorQuestions people actually ask
What is a hallucination escape rate?
The share of responses that reach a user carrying a factual error the system should have caught. It is the hardest of the four objectives to measure and the one worth arguing about — the argument itself usually improves the product.
Which tools do you use?
Whatever you already run, if it is adequate. OpenTelemetry, Prometheus and Grafana are the default because they are open, portable, and nobody has to renew a contract to keep their dashboards.
Can this be part of the audit instead?
The audit tells you what your reliability posture is. This builds it. If you are unsure which you need, start with the audit and credit the fee.
Fifteen minutes is enough to work out whether this is the right engagement, or whether it is one of the others, or none of them.
Book a 15-min callScoping by email works too — send what you're building and the read-back comes with a recommendation on where to start — even when the right start is smaller than you expected.
The SLO starter pack: the four objective definitions, the error budget policy template, and the alerting thresholds behind the 71% to 96% programme.