KV Cache / Multi-head Latent Attention / DeepSeek-V4
PagedAttention assumed MHA. DeepSeek-V4 says that assumption is broken.
Build notes from Tessera: what a KV block manager has to change when the cache is a 56× compressed latent instead of full K/V, what changes again when attention goes per-layer hybrid, and why content-addressed dedup needs byte verification.
What MLA changes about KV cache accounting
Multi-head Latent Attention stores a compressed latent c_kv plus a small RoPE component instead of full per-head keys and values. For DeepSeek-V3's geometry — 61 layers, 128 heads, head dim 128, latent dim 512, RoPE dim 64 — the difference is not incremental.
| Context | MHA BF16 | MLA BF16 | MLA FP8 | Ratio |
|---|---|---|---|---|
| 8K | 30.5 GB | 0.54 GB | 0.30 GB | 56.6× |
| 32K | 122.0 GB | 2.17 GB | 1.21 GB | 56.2× |
| 128K | 488.0 GB | 8.68 GB | 4.84 GB | 56.2× |
| 512K | 1.95 TB | 34.72 GB | 19.4 GB | 56.2× |
| 1M | 3.90 TB | 69.44 GB | 38.7 GB | 56.2× |
The trap is that a serving stack can support MLA models and still throw the compression away. If the framework expands W_UK · c_kv into full K/V before handing blocks to the cache, the block manager is storing MHA-shaped data for an MLA model, and the 56× exists only in the paper. Storing c_kv and k_rope natively — never materialising full K/V at the cache layer — is the whole premise.
Block size is not a tuning knob here
PagedAttention's 16-token block is a good default for MHA, where a block holds a lot of bytes and fragmentation dominates. Under MLA a 16-token block holds roughly one fifty-sixth of that, which turns per-block metadata into a meaningful fraction of the cache and makes the allocator do far more work for the same context.
Tessera uses 64-token blocks for MLA because that is what FlashMLA wants natively. For the V4 hybrid, block size has to be a common multiple of both compression strides — lcm(k1, k2) = 128 for V4-Pro — or a block boundary lands mid-group and the accounting stops being expressible at all.
That is the general shape of the problem: block size is a consequence of the attention geometry, not a knob you tune afterwards.
What DeepSeek-V4 breaks
“The hybrid attention mechanism violates fundamental assumptions behind PagedAttention and its variants.”
V4 does not use one attention mechanism. It interleaves three across layers — compressed sparse attention, hyper-compressed attention, and sliding-window attention — at different compression ratios and different precisions per region. Per-token storage cost stops being a model-level constant and becomes a per-layer one.
| Scheme | Compression | Bytes/token/layer | Composition |
|---|---|---|---|
| CSA | k1 = 4 (overlapping) | 160 B | (64·BF16 + 448·FP8 + 128·FP4) / 4 |
| HCA | k2 = 128 | 4 B | (64·BF16 + 448·FP8) / 128 |
| SWA | uncompressed, window 128 | 576 B | 64·BF16 + 448·FP8 |
Three consequences follow for a block manager. Compression scheme becomes a per-layer map rather than a global enum. Precision becomes per-region *within* a single block — BF16 for RoPE, FP8 for content, FP4 for the indexer — so a single dtype field can no longer describe a block. And a two-tier cache appears: paged blocks for the compressed layers, plus a per-request arena for the sliding-window and uncompressed-tail regions, which do not page well at all.
None of this needs a GPU to get right. The byte accounting is arithmetic over the paper's constants, pinned by unit tests that assert 160, 4 and 576 exactly. If the layout is wrong, the tests fail on a laptop. What a GPU is needed for is the kernel runtime, and that is upstream work — Tessera mounts on FlashMLA, FlashInfer and DeepSeek's TileLang reference implementation; it does not ship a kernel.
Position independence beats prefix matching
Multi-agent pipelines recompute the same context constantly: the same system prompt, the same retrieved documents, the same tool schemas, once per agent. The existing answers to this are prefix caching and cross-model KV transfer, and both are working around a constraint MLA does not have.
- RadixAttention and automatic prefix caching require an exact token-prefix match. Two agents holding the same document behind different preambles share nothing, even though the document's KV is identical.
- KVCOMM-style cross-agent transfer works on any model, but has to *estimate* the KV under a different positional context, and estimation carries error.
- MLA latents are position-independent. The same content produces the same
c_kvwherever it sits in the sequence. So the right key for reuse is a content hash of the latent, not the token prefix — and the reuse is exact, with zero estimation error.
That turns cross-agent sharing into a content-addressed store with reference-counted copy-on-write: hash the block on seal, look it up, share it on match, fork it on write. Sixteen agents over a shared document converge on one copy.
The finding: content-addressed dedup needs byte verification
Content addressing has a failure mode a single-tenant cache never surfaces. The hash used for block dedup was xxh3 — fast, excellent distribution, and not cryptographic. On a hash match, the block manager deduplicated.
In a single-agent cache, a collision is a correctness bug you might never notice. Across agents, a collision means one agent's context is served to another agent's request. That is not a performance regression, it is a context leak — and under a non-cryptographic hash it does not have to happen by accident. It requires an adversary who can influence content and compute a preimage.
The fix is defence in depth, not a bigger hash. On a hash match, compare the bytes; if they differ, install the candidate as a fresh block and increment a dedup-collision counter. The counter is the operator-facing half: a non-zero value there is a security signal, not a tuning hint. Migrating to a cryptographic hash stays on the table, but byte verification is what makes the property structural rather than probabilistic.
The same audit produced a second change worth generalising. Device pointers were raw machine addresses carried next to a length. Two blocks could alias and nothing in the type system objected. Replacing them with { region, offset, len } handles — where the region is an index into a table the backend owns — makes aliasing not merely unlikely but unrepresentable, and turns pointer resolution into O(1) indexing as a side effect. The pattern generalises: when a safety property depends on discipline, move it into a type.
What is verified, and what is not
Everything above runs on a CPU-only machine. Block layout, seal-and-dedup, copy-on-write, tiered eviction, the per-request state cache, the filesystem-backed disk tier with its three sliding-window persistence strategies, quarantine-on-checksum-mismatch recovery, and the paper-constant byte accounting are all CPU-validated, with 24 ADRs recording the decisions behind them.
GPU-gated and therefore not claimed: FlashMLA and FlashInfer parity against a PyTorch reference oracle, a 128K needle-in-haystack precision regression, throughput against stock vLLM, and integration against a live vLLM engine. The V4 kernel runtime itself is upstream-pending. There is no number in this note that a GPU would be required to produce.
Tessera is a proof of concept built in public, dual-licensed MIT or Apache-2.0. The compression ratios here are structural — arithmetic over documented model geometry, verified by tests — not measurements of a running system.
Key takeaways
- Multi-head Latent Attention compresses DeepSeek-V3's KV cache about 56× against MHA at BF16 — 488 GB down to 8.68 GB at 128K context. Most serving stacks give that back by expanding the latent into full K/V before caching it.
- A 16-token block is correct for MHA and wrong by two orders of magnitude for MLA. Tessera uses 64-token blocks for MLA and lcm(k1, k2) = 128 for the V4 hybrid.
- DeepSeek-V4 makes compression a per-layer property: CSA at 160 bytes per token per layer, HCA at 4, SWA uncompressed at 576. A single global scheme cannot express that model.
- MLA latents are position-independent, so identical context across agents is literally the same content. Content-addressed hashing gives exact cross-agent reuse with no prefix matching and no estimation error.
- Content-addressed dedup under a non-cryptographic hash is a context-leak surface. Verify bytes on hash match, and treat collisions as a security metric.
Frequently asked questions
What is Multi-head Latent Attention (MLA)?
An attention variant used by DeepSeek-V3 and Kimi-K2 that caches a low-rank compressed latent plus a small RoPE component instead of full per-head keys and values. It reduces KV cache size roughly 56× versus MHA at BF16, and the latent is position-independent, which is what makes content-addressed reuse across requests possible.
Does Tessera replace FlashMLA or FlashInfer?
No. Tessera is the block-layout and accounting layer underneath them. Attention kernels stay upstream — FlashMLA on SM 9.0 and above, FlashInfer on Ampere and newer, TileLang for the V4 hybrid, Triton as fallback — and Tessera mounts on whichever the dispatcher selects.
Why not just use vLLM's prefix caching?
Prefix caching keys on an exact token prefix. Two agents that share a document but differ in their preamble share nothing. MLA latents are position-independent, so hashing the latent content lets those agents share the document's KV directly, regardless of where it appears in each sequence.
Is the 56× compression number measured or calculated?
Calculated, and verified structurally. It is arithmetic over DeepSeek-V3's published geometry, pinned by unit tests asserting the per-token byte counts. It is not a throughput measurement, and no throughput claim is made anywhere in the project.
Is Tessera production ready?
No. It is a proof of concept. It is CPU-validated with CI, 24 ADRs and a security-hardening pass behind it; the GPU paths, the V4 kernel runtime and multi-node transport are wired but unvalidated.
Sources & references
Related notes
- Reasoning models emit two workloads. Your scheduler sees one.
Build notes from Meridian: why a reasoning request should be scheduled as two workloads with separate SLOs, how entropy replaces the token timer, and what broke when I ran the toolchain instead of reading the code.
- How to publish infrastructure work you can't fully benchmark
Most infrastructure side projects are written on a laptop and benchmarked on a slide. A per-component verification column is cheaper than a bigger claim, and worth more.
Updates on new engineering notes: production AI infrastructure, with the arithmetic shown.
Request by emailn@nickcerutti.com · Include the question or project you have in mind.
Apply this to your system
Cache behavior belongs in the platform design, alongside tenant boundaries, routing and the cost of serving the workload.