Agent Production Audit
Agents are in production and nobody can say whether they are safe.
Two weeks: an independent read of cost, reliability, and governance across your agent stack, with every finding priced in dollars.
2 weeks · from $12K · Signed off by CTO or VP Engineering
1.8M+
Tool invocations audited under deterministic policy validation
Zero
Known policy bypasses across audited deployments
250-450
Concurrent agent instances operated in production
Who this is for
- Agents are live, the team has grown, and nobody has read the whole system in one pass.
- You need an outside opinion before a board conversation, a funding round, or a security review.
- You suspect the problems are real but you cannot rank them.
And who it is not
- You want automated scanning on its own. Scanners are excellent at detecting; this engagement is the judgment layer — what the findings mean for your architecture and your bill, ranked by dollar impact.
- You want a compliance certificate. That is a different engagement, and it is on the compliance pages.
What you get
- A written audit across multi-tenancy and isolation, spend control, reliability targets, retrieval, and the tool-call layer.
- Findings ranked by dollar impact, with the arithmetic shown for each.
- A remediation plan split into what your team can do this month and what needs a dedicated engagement.
- An Agent Bill of Materials: what agents exist, what they can reach, and under whose authority.
- A walkthrough with your engineering team, plus a version of the summary that survives being forwarded to a non-engineer.
How it runs
Week 1
Read the system
Architecture, traces, policies, and the deployment path. Interviews with the people who operate it, because the documented system and the running system are rarely the same one.
Week 2
Rank and price
Every finding gets a dollar figure and a fix. The report is delivered and walked through, not emailed and abandoned.
This has shipped
01
Production emergency
$40K+/day burn contained within 72 hours. Zero recurrence.
03
SLO framework program
SLO compliance from 71% to 96% over two quarters.
04
Retrieval and governance at scale
1.8M+ tool invocations under policy validation. Zero known bypasses.
05
LLM cost reduction program
$200K/yr aggregate reduction in cloud and LLM spend.
07
Real-time trading system
Sub-10ms latency at peak. Slippage cut 14%, zero limit breaches on real capital.
Read the code first
What it costs
from $12K
Scales with the number of services, models, and tenants in scope. Set on the call.
Every engagement can start as a $2,500 paid pilot, credited in full against the full scope if you go ahead.
Every engagement and what it costsBefore you talk to anyone
Agent SLO & Error Budget Calculator
Turn reliability targets into error budgets, alert thresholds, and a number your team can argue about before the incident. Free, runs in your browser, and nothing you type leaves it.
Open the agent slo & error budget calculatorQuestions people actually ask
What makes this different from running the free scanners?
The scanners detect. They do not decide. The output of a scan is a list of findings with no idea which one costs you $9K a month and which one is noise. Ranking that list against your architecture and your bill is the deliverable.
Will you sign off on our security posture?
The audit gives you an evidence-backed opinion; a formal attestation comes from an accredited auditor, and independence is what gives it value — the two stay separate by design.
Can the audit fee be credited?
Yes. If the audit leads directly into a cost sprint, a reliability engagement, or a governance package, the fee is credited against it.
How much of our team's time does this take?
Roughly four hours in week one for access and interviews, and two hours in week two for the walkthrough.
Fifteen minutes is enough to work out whether this is the right engagement, or whether it is one of the others, or none of them.
Book a 15-min callScoping by email works too — send what you're building and the read-back comes with a recommendation on where to start — even when the right start is smaller than you expected.
The SLO starter pack: the four objective definitions, the error budget policy template, and the alerting thresholds behind the 71% to 96% programme.