📄 Free guide · 26 pages

The KV Cache, from zero.

Every AI company is now spending real money on something most people have never heard of. This explains what it is, why it ends up bigger than the model itself, and why the entire memory and storage industry is rebuilding around it — starting from nothing.

No maths. No jargon that isn't defined on the page where it appears.

Free · PDF

The KV Cache — from zero

26 pages · about a 25-minute read
  • What a layer actually is, and why a model stacks 80 of them.
  • What Q, K and V mean — as a meeting where everyone wears a name tag.
  • Why two of them are kept and one is thrown away. That lopsided fact is the whole cache.
  • The actual arithmetic: 320 KB per word, 40 GB per conversation, and where every number comes from.
  • Why a busy server holds more conversation than model — and what that does to the price of every word.
  • Why storage companies suddenly care, and what they shipped about it.
Cover page of the KV cache explainer A model as an assembly line of layers Diagram of work repeated without a cache The memory ladder: desk, shelf, filing room, warehouse
Written by Nathan Wang (AI-Nate)ai-nate.com

Want to build the things this describes?

The guide explains how the machinery works. The cohort is where you build on top of it — live, hands-on, with the systems running by the end.

See the cohort