Skip to content
All articles

Agentic AI · Context

Your context window is a budget

Every token in the window costs money, latency and the model's attention. Reserve room for the answer first, then spend what is left on whole bundles of evidence, never on half of one.

· 5 min read

A token budget for a hypothetical 8,192-token window: 1,024 reserved for the answer and 512 for message framing and estimation error leave an input allowance of 6,656. Instructions take 700, tool schemas 1,000, and the current task plus typed state 700, leaving 4,256 for evidence and history. Six 800-token evidence chunks need 4,800 and are over by 544; five need 4,000 and fit, but may drop the exception. Rank whole bundles: a rule plus its exceptions, all or nothing.

Every token you put in front of a model is paid for three times: in money, in latency, and in the model's attention. A bigger context window raises the ceiling. It does not change the job, which on every turn is to send the smallest set of tokens that lets the model answer correctly.

Five lines in the budget

Budget lineGrows withLever
System promptPolicy and operating instructionsKeep it precise, versioned, and stable where caching applies
Tool schemasThe capabilities exposed in this stateExpose tools per task, or let the agent discover them progressively
Conversation stateRaw turns, summaries, persisted factsCompact it, and retrieve what this decision needs
Retrieved contextNumber of chunks × chunk sizeRerank so the count stays small and precise
Long-term memoryHow much is pulled in per turnRetrieve the few relevant memories, not the whole profile

Two of these lines are easy to underestimate. In runtimes that send every registered tool on every turn, irrelevant schemas use up input tokens and can make choosing the right tool harder. And replaying every raw turn of a conversation grows without limit unless the application compacts it, summarizes it, or retrieves only the state the current decision needs.

Do the arithmetic before choosing context

Take a hypothetical 8,192-token window. Reserve 1,024 tokens for the answer and 512 for message framing and estimation error, which leaves an input allowance of 6,656. Instructions take 700, tool schemas 1,000, the current task 200 and typed state 500: 2,400 spent before any evidence, and 4,256 left for evidence and history.

8,192 - 1,024 answer - 512 framing      = 6,656 input allowance
6,656 - (700 + 1,000 + 200 + 500)       = 4,256 for evidence and history

6 chunks x 800 = 4,800   over by 544
5 chunks x 800 = 4,000   fits, with 256 to spare

Use your provider's exact token accounting, including any reasoning or output limits. A rough counter is only an estimate, and the budget is only as good as its numbers.

Common Mistake

Taking the top five chunks because five fit. The sixth may be the exception to a rule stated in the first five, and a prompt that holds the rule without its exception produces a confident, wrong answer. Rank whole evidence bundles, a rule together with its exceptions, and include each bundle completely or not at all. If the bundle the question needs does not fit, report insufficient context rather than answering from part of it.

The principle that makes retrieval work, a handful of highly relevant chunks beating a large dump, applies to the whole session. Every line of the budget competes for the same window.

Get one diagram a week

A short article built around one engineering diagram, from the same library as these courses.

One diagram-led article a week on AI and systems engineering. We email you once to confirm, and every newsletter has an unsubscribe link. Privacy policy