
Aki Wijesundara
Manu Jayawardana
The number on the model card. The number in the spec sheet. The number your CTO put in a quote.
Much lower. Different per task. Never the same as the advertised number. This is the gap where teams get burned.


Every token attends to every other token. That's your baseline attention budget.
Double the context and you quadruple the pairwise relationships. Same fixed budget. Thinner attention on each pair. Not smarter. Just spread.

Retrieval accuracy is high at the start and end of the window, and sags in the middle. Same fact, different position, different answer.

Higher chance of the right answer. No ambiguity about which document to believe. Nothing competing with the truth.
The model can't tell which is authoritative. Accuracy drops. This is why "top-k retrieval" without conflict handling gets worse, not better, as k rises.


Single-shot QA is one big blob. An agent accumulates its own exhaust.

Same prompt. Same question. Adding more retrieved material makes the response softer, less specific, less useful.
The model quotes a plausible but incorrect passage. It's not making things up. It's being misled by a distractor in the window.
You didn't change the prompt. You changed the surroundings. The instruction is drowning in nearby noise.
Each request keeps loading more into the window. Bills climb. Answers stay flat or degrade. That's a distraction tax.
Not the most context. Not the biggest window. The smallest high-signal set. Every token you add costs attention. Every token you remove has to actually be dead. This is the whole discipline in one sentence.


config.v1.json config.final.json config.FINAL2.json
The model will pick the wrong one with confidence. It has no way to know which is current, and the file names actively lie to it.
Everything else is gone or archived outside the window. The model can only cite the version that's actually current.

Say it out loud. Apply it to every block of context, one at a time. Decoration goes. The exam is: would the answer change?

Full chat history, full tool output, full recall of prior conversations. All of it is decoration disguised as memory.
The next three passes give you the moves that keep the memory but drop the token weight.



You pay tokens for every tool description whether the agent uses it or not. And tool selection accuracy degrades as the menu grows. More options, worse choices.
Planning and execution should not share the same surface. Give the planner one set. Give each executor a scoped, minimal set. Selection accuracy jumps. Token spend drops.



The window fills with maybes. Half of them go unused. All of them steal attention from the ones the agent actually reaches for.
Navigation beats speculation. The agent asks for what it needs, when it needs it. The window stays tight. The reasoning stays fresh.

Summarise and reinitialise. Keep decisions and constraints. Discard redundant tool output. Fastest, cheapest, most disruptive to cache.
In-placeWrite state to a file outside the window. Pull it back on demand across thousands of steps. Cache-friendly. Cheap on tokens; costs a filesystem.
ExternalA specialist burns 10k tokens in a clean window and returns a 1.5k summary. Real gain when branches are independent. Coordination overhead when they aren't.
Parallel


The window refills every turn, and nobody is steering it unless you build the harness. Context engineering as a first-class skill, not a tip. Retrieval, memory, evals, and agent loops you can measure. Shipped systems, not slideware.
Nine weeks. Context engineering, evals, memory, and production agent loops. Certificate for engineers and AI PMs.
Where this leadsPost one line: the biggest context block you're afraid to trim. Aki or Manu will read a few aloud and diagnose live.
Open floorInstrument one agent. Log tokens per turn. Delete the worst duplicate. Send us the before/after. We reply with a short audio review.
The receipt