Skip to content

Nick Cerutti

AI Architect & Consultant · Production LLMs & Agents

I design and harden the systems that break after the demo. Agents, inference, and multi-tenant platforms that survive contact with real users and real budgets.

Portrait of Nick Cerutti

Open to new engagements

Book a 15-min call

Experience

Professional experience

Systems, infrastructure, and advisory work across trading, production AI, and distributed engineering organizations.

Client work

Production work, with the numbers left in.

Eight client work examples — seven anonymized, one public (Iguazio). Every metric below also appears in the experience section above.

01

Production emergency

$40K+/day burn contained within 72 hours. Zero recurrence.

02

Multi-tenant agent substrate

250-450 concurrent agent instances in production. Release overhead down roughly 40%.

05

LLM cost reduction program

$200K/yr aggregate reduction in cloud and LLM spend.

Projects

Selected open-source work

Tooling for agent memory, governance, orchestration, and the operational realities around production deployment — built in public at github.com/angelnicolasc.

Drop-in persistent memory for AI agents. Cut token consumption by 97% with a single zero-dependency binary. No Docker, no external databases, no cloud storage. Pure pragmatism. +450 GitHub stars and 35 forks.

Golang / Agent Memory / Hybrid Retrieval / MCP Server / Persistent State

Embeddable Rust library for formal causal belief revision in LLM agents. Grounded in 5 arXiv papers, it delivers dual-process uncertainty quantification and cryptographically auditable trails.

Rust / PyO3 / Causal Graphs / Agent Governance / MCP

Inference-time compute scheduler for reasoning models, optimizing token-generation phases via dual-queue dispatch and entropy-driven budget control.

Rust / Python / CUDA / Inference Infrastructure / vLLM

Local dual-tier LLM serving stack featuring sub-millisecond adaptive complexity routing and real-time VRAM orchestration.

Python / vLLM / llama.cpp / Infrastructure Optimization / ASGI

Open-core security and governance platform for production AI agents, spanning static analysis, policy authoring, and deterministic runtime enforcement.

Python / Policy-as-Code / Runtime Governance / AI Agent Security

Universal harness that wraps existing agent systems to add observability, memory, budget controls, and safe self-evolution without forcing framework migration.

Python / Multi-Agent Orchestration / Self-Optimizing Runtime / Observability

Work with me

Production-minded engagement.

Three common starting points. Nine consulting services in total.
Start from what is actually broken — each of these opens the engagement built for it. The full list is one click below.

Spend is growing faster than usage.

LLM Cost & Routing Sprint

3 weeks · $15-20K

Agents are in production and nobody can say whether they are safe.

Agent Production Audit

2 weeks · from $12K

The team is shipping AI and nobody owns the architecture.

Fractional AI Architect

3 months minimum · $9-15K/mo

Agents burning money right now? Email me directly — skip the calendar.

If you're building something where AI has to work in production, let's talk.

Book a 15-min call

No deck, no pitch. Not sure which one fits? Send me what you're building and you'll get a straight read back — scope, sequence, and where a $2,500 pilot makes sense before anything bigger.

On the employment track, I'm also open to full-time Head of AI and AI Architect roles. Download the resume or ask in the call.

Notes

Engineering notes

Short technical notes from real infrastructure work. The kind of details that usually stay inside private Slack channels or incident reviews. Last updated .

Itemising what one architectural component of a 284B model actually costs to serve - and why the published description of it was wrong by 2x.

Inference / Cost Modeling / Open Source / DeepSeek-V4

6 min read

Most infrastructure side projects are written on a laptop and benchmarked on a slide. A per-component verification column is cheaper than a bigger claim, and worth more.

Engineering Practice / Open Source / Benchmarking / Architecture Decision Records

8 min read

Capabilities

Technical focus

Broad enough to span research, infra, and runtime operations.
Specific enough to ship in production.

Languages & Runtimes

Core implementation languages for systems work, infrastructure tooling, and performance-sensitive runtime paths.

Go, Python, TypeScript, Rust, SQL, C++, CUDA.

Agent Systems

For orchestrating multi-step agent behavior, memory, tool use, and compound execution across frameworks.

Multi-agent orchestration, A2A protocol, MCP server and client development, Context engineering, Persistent and episodic memory systems, Hybrid retrieval-augmented generation (RAG), Tool-use and function-calling pipelines.

Inference & Serving

From single-GPU deployments to multi-tenant inference clusters under real production load.

VLLM, Triton Inference Server, Model routing, AWQ and GPTQ quantization, KV-cache optimization, Prompt caching, Structured output contracts, Low-Rank Adaptation (LoRA) and Quantized Low-Rank Adaptation (QLoRA) pipelines.

Scale & Reliability

For systems that need to stay observable, cost-aware, and resilient as traffic, tenants, and model complexity grow.

Multi-tenant agent platforms, Kubernetes at scale, Service Level Objective (SLO) and error budget design, Incident response, Chaos engineering, Capacity planning, OpenTelemetry, Prometheus and Grafana, Agent trace instrumentation, Real-time memory analytics, Token spend tracking, Audit trails, Evaluation loops.

Edge & Constrained Environments

When Kubernetes is too much and reliability still is not optional.

Zero-dependency binaries, Embedded storage, Offline-first agents, Single-binary MCP servers, Bubble Tea observability, WASM-compatible tooling.

Security & Governance

Enforcement that runs at runtime, not just on paper, across agent policies, compliance, and tool-call controls.

Policy-as-code, Runtime tool-call validation, Deterministic evaluation loops, Reachability analysis, Agent BOM, Desktop agent scanning, Prompt injection mitigation.

Delivery & Automation

For shipping changes repeatedly, enforcing standards, and reducing manual operational overhead.

GitHub Actions, ArgoCD, MCP integrations, CLI tooling, Pre-commit hooks, SARIF, GitOps pipelines.

Education

Academic background

Computer science, business, and applied ML training with an emphasis on systems thinking.

S21 University logo

S21 University

B.S. in Computer Science

GPA 3.92

Buenos Aires, Argentina

Coursework

Data Structures and Algorithms, Advanced Operating Systems, Software Engineering, Database Systems, Computer Architecture, Entrepreneurial Development.

Jan. 2018 - May 2021
IAE Business School logo

IAE Business School

Master of Business Administration (in progress)

Buenos Aires, Argentina / Online

Coursework

Corporate Finance, Strategy & Business Model Innovation, Financial Accounting, Operations & Logistics, Leadership & Human Behavior.

Feb. 2025 - May 2027
DeepLearning.AI / Coursera logo

DeepLearning.AI / Coursera

MLOps Specialization

Online

Coursework

Introduction to Machine Learning in Production, Machine Learning Data Lifecycle in Production, Machine Learning Modeling Pipelines in Production, Deploying Machine Learning Models in Production.

Mar. 2024