- Brought in across 6+ fast-growing platforms, from Series B to late-stage teams, to build the multi-tenant substrate for agent systems serving 250-450 concurrent instances across AWS Bedrock, AgentCore, and Azure environments.
- Led incident response for a production multi-agent cost explosion projected at $40K+ per day, implementing circuit breakers, per-session spend caps, and retrospective governance controls within 72 hours with zero recurrence.
- Defined Service Level Objective (SLO) frameworks across 3 client platforms using task completion rate, mean tool calls, hallucination escape rate, and end-to-end latency; improved SLO compliance from 71% to 96% over two quarters.
- Audited and built retrieval and governance systems supporting 1.8M+ tool invocations, with deterministic policy validation and zero known bypasses across deployments using Pinecone and OpenSearch.
- Built Go, Python, and TypeScript middleware, Azure DevOps templates, and deployment automation for 4 distributed engineering teams, reducing release-cycle overhead by roughly 40%.
- Refactored client integrations with Anthropic SDK, n8n, and Copilot Studio to enforce token optimization, caching, and low-infrastructure patterns, delivering an aggregate $200K annual reduction in cloud and LLM spend.
- Standardized data lakes and model registries across 2 client platforms, integrating Databricks and Domino Data Lab with classical ML workloads in scikit-learn, PyTorch, and TensorFlow.
Nick Cerutti
AI Architect & Consultant · Production LLMs & Agents
I design and harden the systems that break after the demo. Agents, inference, and multi-tenant platforms that survive contact with real users and real budgets.

Open to new engagements
Book a 15-min callExperience
Professional experience
Systems, infrastructure, and advisory work across trading, production AI, and distributed engineering organizations.
Client work
Production work, with the numbers left in.
Eight client work examples — seven anonymized, one public (Iguazio). Every metric below also appears in the experience section above.
Projects
Selected open-source work
Tooling for agent memory, governance, orchestration, and the operational realities around production deployment — built in public at github.com/angelnicolasc.
Drop-in persistent memory for AI agents. Cut token consumption by 97% with a single zero-dependency binary. No Docker, no external databases, no cloud storage. Pure pragmatism. +450 GitHub stars and 35 forks.
Golang / Agent Memory / Hybrid Retrieval / MCP Server / Persistent State
02
Epica
Embeddable Rust library for formal causal belief revision in LLM agents. Grounded in 5 arXiv papers, it delivers dual-process uncertainty quantification and cryptographically auditable trails.
Rust / PyO3 / Causal Graphs / Agent Governance / MCP
03
Meridian
Inference-time compute scheduler for reasoning models, optimizing token-generation phases via dual-queue dispatch and entropy-driven budget control.
Rust / Python / CUDA / Inference Infrastructure / vLLM
04
Stratum
Local dual-tier LLM serving stack featuring sub-millisecond adaptive complexity routing and real-time VRAM orchestration.
Python / vLLM / llama.cpp / Infrastructure Optimization / ASGI
05
Drako
Open-core security and governance platform for production AI agents, spanning static analysis, policy authoring, and deterministic runtime enforcement.
Python / Policy-as-Code / Runtime Governance / AI Agent Security
06
Forge
Universal harness that wraps existing agent systems to add observability, memory, budget controls, and safe self-evolution without forcing framework migration.
Python / Multi-Agent Orchestration / Self-Optimizing Runtime / Observability
Work with me
Production-minded engagement.
Three common starting points. Nine consulting services in total.
Start from what is actually broken — each of these opens the engagement built for it. The full list is one click below.
Spend is growing faster than usage.
LLM Cost & Routing Sprint
3 weeks · $15-20K
Agents are in production and nobody can say whether they are safe.
Agent Production Audit
2 weeks · from $12K
The team is shipping AI and nobody owns the architecture.
Fractional AI Architect
3 months minimum · $9-15K/mo
Agents burning money right now? Email me directly — skip the calendar.
If you're building something where AI has to work in production, let's talk.
Book a 15-min callNo deck, no pitch. Not sure which one fits? Send me what you're building and you'll get a straight read back — scope, sequence, and where a $2,500 pilot makes sense before anything bigger.
On the employment track, I'm also open to full-time Head of AI and AI Architect roles. Download the resume or ask in the call.
Notes
Engineering notes
Short technical notes from real infrastructure work. The kind of details that usually stay inside private Slack channels or incident reviews. Last updated .
Build notes from KVSentry: why per-tenant cache isolation is not the zero-cost mitigation it gets quoted as, how to price it from lost cache hit rate instead of a published percentage, and the one metric that moves the answer by two orders of magnitude.
FinOps / KV Cache / Multi-Tenant Security / vLLM
9 min read
Itemising what one architectural component of a 284B model actually costs to serve - and why the published description of it was wrong by 2x.
Inference / Cost Modeling / Open Source / DeepSeek-V4
6 min read
Most infrastructure side projects are written on a laptop and benchmarked on a slide. A per-component verification column is cheaper than a bigger claim, and worth more.
Engineering Practice / Open Source / Benchmarking / Architecture Decision Records
8 min read
Capabilities
Technical focus
Broad enough to span research, infra, and runtime operations.
Specific enough to ship in production.
Languages & Runtimes
Core implementation languages for systems work, infrastructure tooling, and performance-sensitive runtime paths.
Go, Python, TypeScript, Rust, SQL, C++, CUDA.
Agent Systems
For orchestrating multi-step agent behavior, memory, tool use, and compound execution across frameworks.
Multi-agent orchestration, A2A protocol, MCP server and client development, Context engineering, Persistent and episodic memory systems, Hybrid retrieval-augmented generation (RAG), Tool-use and function-calling pipelines.
Inference & Serving
From single-GPU deployments to multi-tenant inference clusters under real production load.
VLLM, Triton Inference Server, Model routing, AWQ and GPTQ quantization, KV-cache optimization, Prompt caching, Structured output contracts, Low-Rank Adaptation (LoRA) and Quantized Low-Rank Adaptation (QLoRA) pipelines.
Scale & Reliability
For systems that need to stay observable, cost-aware, and resilient as traffic, tenants, and model complexity grow.
Multi-tenant agent platforms, Kubernetes at scale, Service Level Objective (SLO) and error budget design, Incident response, Chaos engineering, Capacity planning, OpenTelemetry, Prometheus and Grafana, Agent trace instrumentation, Real-time memory analytics, Token spend tracking, Audit trails, Evaluation loops.
Edge & Constrained Environments
When Kubernetes is too much and reliability still is not optional.
Zero-dependency binaries, Embedded storage, Offline-first agents, Single-binary MCP servers, Bubble Tea observability, WASM-compatible tooling.
Security & Governance
Enforcement that runs at runtime, not just on paper, across agent policies, compliance, and tool-call controls.
Policy-as-code, Runtime tool-call validation, Deterministic evaluation loops, Reachability analysis, Agent BOM, Desktop agent scanning, Prompt injection mitigation.
Delivery & Automation
For shipping changes repeatedly, enforcing standards, and reducing manual operational overhead.
GitHub Actions, ArgoCD, MCP integrations, CLI tooling, Pre-commit hooks, SARIF, GitOps pipelines.
Education
Academic background
Computer science, business, and applied ML training with an emphasis on systems thinking.

S21 University
B.S. in Computer Science
GPA 3.92
Buenos Aires, Argentina
Coursework
Data Structures and Algorithms, Advanced Operating Systems, Software Engineering, Database Systems, Computer Architecture, Entrepreneurial Development.
Jan. 2018 - May 2021

IAE Business School
Master of Business Administration (in progress)
Buenos Aires, Argentina / Online
Coursework
Corporate Finance, Strategy & Business Model Innovation, Financial Accounting, Operations & Logistics, Leadership & Human Behavior.
Feb. 2025 - May 2027

DeepLearning.AI / Coursera
MLOps Specialization
Online
Coursework
Introduction to Machine Learning in Production, Machine Learning Data Lifecycle in Production, Machine Learning Modeling Pipelines in Production, Deploying Machine Learning Models in Production.
Mar. 2024