# Nick Cerutti > Independent AI architecture consultant for production LLMs and agents. Remote consulting for engineering teams: cost optimization, production reviews and fractional architecture. The linked pages are the source of scope, pricing, evidence and limitations. Client work includes anonymized engagements and one public employer engagement; results are not guarantees for a new project. Regulatory tools describe technical readiness, not legal advice. ## Start here - [Profile and contact](https://nickcerutti.com/): Background, experience and contact options. - [Pricing and terms](https://nickcerutti.com/pricing): Published prices, engagement terms and scope. ## Services - [LLM Cost & Routing Sprint](https://nickcerutti.com/read/services/llm-cost-sprint.md): Reduce LLM spend with routing, prompt caching and runtime budgets. Three-week consulting sprint with clear deliverables. From $15K. - [Agent Production Audit](https://nickcerutti.com/read/services/agent-production-audit.md): Independent review of agent cost, reliability and governance. A two-week audit with prioritized findings and a remediation plan. From $12K. - [Agent Governance & Insurability Package](https://nickcerutti.com/read/services/agent-insurability.md): Prepare agent controls and evidence for insurance and security reviews: inventory, identity, least privilege and audit trails. Three to four weeks. - [AI Compliance Package — EU Article 50 plus US module](https://nickcerutti.com/read/services/ai-act-article-50.md): Implement transparency and content-marking controls for production AI systems. Three-week technical engagement alongside counsel. From $15K. - [Fractional AI Architect](https://nickcerutti.com/read/services/fractional-ai-architect.md): Retained architecture support for production LLMs and agents: system design, evaluations, reliability and cost. From $9K/month; three-month minimum. - [Reliability Foundation — SLOs and evals in CI](https://nickcerutti.com/read/services/reliability-foundation.md): Define agent SLOs, error budgets and evaluations in CI. Three-week consulting engagement for production reliability. From $12K. - [Substrate Build](https://nickcerutti.com/read/services/substrate-build.md): Build tenant isolation, routing, memory, evaluations and spend controls for production agents. Six to eight weeks, with handover to your team. - [Pre-audit preparation — LL144 and bias audits](https://nickcerutti.com/read/services/bias-audit-prep.md): Prepare hiring AI for an independent bias audit: technical gap analysis and remediation. One to two weeks. This is preparation, not the formal audit. - [Production emergency](https://nickcerutti.com/read/services/production-emergency.md): Contain runaway agent spend and production failures with triage, cost attribution, circuit breakers and spend caps. Contact Nick directly for scoping. ## Client work - [All client work](https://nickcerutti.com/work): Eight engagement summaries: seven anonymized, one public. - [Containing Runaway AI Agent Costs](https://nickcerutti.com/read/work/emergency-40k.md): An anonymized incident-response engagement: projected $40K/day agent spend contained within 72 hours through attribution, circuit breakers and session budgets. - [Building Multi-Tenant AI Agent Platforms](https://nickcerutti.com/read/work/substrate.md): Production agent platform work across client teams: tenancy, routing, memory, evaluation and cost controls, followed by deployment and engineering handover. - [LLM Cost Optimization Across Client Platforms](https://nickcerutti.com/read/work/cost-program.md): A portfolio-wide cost program addressing repeated context, prompt caching and token-heavy integrations. $200K/year aggregate reduction in cloud and LLM spend. ## Engineering notes - [The Cost of KV Cache Isolation](https://nickcerutti.com/read/notes/pricing-the-security-fix-nobody-prices.md): Price per-tenant KV cache isolation through lost cache hits and extra computation. KVSentry build notes with assumptions and validation limits. - [The Hidden Costs of DeepSeek-V4 Inference](https://nickcerutti.com/read/notes/the-cost-you-cant-see.md): A source-code-based cost model for one DeepSeek-V4 component, including the gap between its published description and implementation. - [Publishing Infrastructure Without GPU Benchmarks](https://nickcerutti.com/read/notes/publishing-infrastructure-you-cannot-benchmark.md): A practical verification matrix for infrastructure projects: distinguish tests, structural checks and unmeasured performance without inventing results. - [KV Cache Layout After MLA](https://nickcerutti.com/read/notes/kv-cache-block-layout-after-mla.md): How compressed latents, hybrid attention and cross-agent reuse change KV block management. Tessera build notes with sources and verification limits. - [Phase-Aware Scheduling for Reasoning Models](https://nickcerutti.com/read/notes/phase-aware-scheduling-for-reasoning-models.md): How Meridian separates think-decode and output-decode workloads, uses entropy signals and handles verification limits. First-hand engineering notes. ## Free tools - [Agent Burn Calculator](https://nickcerutti.com/tools/agent-burn-calculator): What your agents cost per day, and what a runaway loop would cost before anyone noticed. - [AI Regulation Deadline Tracker](https://nickcerutti.com/tools/ai-compliance-deadlines): Every dated AI obligation that applies to a product team, EU and US, with the primary source for each. - [Agent Insurability Scorecard](https://nickcerutti.com/tools/agent-insurability-scorecard): Twenty questions a cyber underwriter or an enterprise security review will ask about your agents. - [EU AI Act Article 50 Self-Assessment](https://nickcerutti.com/tools/ai-act-article-50-check): Which transparency obligations apply to your system, who owes them, and from which date. - [Agent SLO & Error Budget Calculator](https://nickcerutti.com/tools/agent-slo-calculator): Turn reliability targets into error budgets, alert thresholds, and a number your team can argue about before the incident. - [LLM Routing & Cache Savings Estimator](https://nickcerutti.com/tools/llm-routing-estimator): What routing by task class and fixing context re-injection would save you, before anyone writes code. - [MCP Server Security Checklist](https://nickcerutti.com/tools/mcp-security-checklist): Paste an MCP server config and see what a policy engine would flag. Nothing leaves your browser. - [Agent Bill of Materials Generator](https://nickcerutti.com/tools/agent-bom): Build the agent inventory an auditor, an underwriter, or a security reviewer will ask you for. ## Optional - [RSS feed](https://nickcerutti.com/feed.xml): Published engineering notes. - [Privacy](https://nickcerutti.com/privacy): Email requests, attribution and cookieless analytics.