Skip to content
Handbook navigation

Full-access lesson

Context Windows

> Under this chapter's 32-layer, 32-KV-head, 128-dimensional, fp16 MHA configuration, 32 fully occupied 8,192-token sequences require 128 GiB of raw KV payload. GQA, MQA, lower-precision caches, paging, fragmentation, metadata, and workspace change actual allocation.

The complete curriculum

All 12 volumes, companion resources, interviews, and architecture reviews.

Executable engineering practice

Subscriber-only Python, Java, TypeScript, and SQL labs in the isolated runner.

Context Windows | KnowledgeOS