context-budget
Context budget
The context window is RAM, not a hard drive. Full ≠ free: a window stuffed to 95% does not just cost more money, it reasons worse. Model performance degrades as input tokens grow — even well inside the stated limit, every token added depletes a finite attention budget (Chroma "Context Rot" research, accessed 2026-06-02). Your job on a long task is to keep the live window lean and externalise everything else, so the work can run for hours across many fresh windows without losing the thread.
The one rule: if you can reconstruct a thing from a file or from git, it does not belong resident in the window. Keep load-bearing-right-now; evict the rest. The cost of forgetting is one re-read; the cost of hoarding is silent quality rot on every turn that follows.
Neighbours, so you don't do their job here: pricing tokens, spend ledgers and hard $ caps are ../cost-tracking/SKILL.md — same words ("token budget"), different unit, dollars vs. attention. Finding the right context via embeddings/chunking is ../rag/SKILL.md; RAG is how you find context, this is how much you let live and when to evict. Prompt text, few-shot and output format are ../prompt-engineering/SKILL.md; the agent loop, tool schemas and provider adapters are ../building-agents/SKILL.md; partition-then-gather fan-out of independent work is ../parallel/SKILL.md (this skill uses subagents as a context-isolation tactic but does not own that discipline); the 01-TOOLS / 02-DOCS control plane is ../harness/SKILL.md.
Read the gauge first
Before you do anything, estimate utilisation: live input tokens ÷ the model's window limit. You cannot manage a budget you are not watching.
- Compact early — around ~60% utilisation, not 80–95%. Most people only act when quality already broke at 80–95%; by then the rot already happened. Treat 60% as the line where you start reducing, not panicking (practitioner guidance on Claude Code
/compact, accessed 2026-06-02). - Trust the symptoms as an earlier trigger than the number. You can feel rot before the gauge confirms it:
- You re-read a file you already read this session.
- You restate the plan or a decision you already made.
- You contradict an earlier choice.
- Tool results from ten turns ago are still sitting verbatim in the window.