TAI Labs
TAI Lightning Lesson · Free

High-signal
context

The smallest set of tokens. The best chance of the right answer.
Session
Live · one sitting
You bring
One agent whose context has grown out of control
Led by
Aki WijesundaraAki Wijesundara
Manu JayawardanaManu Jayawardana
TAI Labs
The claim

A 1M token window does not mean 1M usable tokens.

Advertised length

What the API will accept without erroring.

The number on the model card. The number in the spec sheet. The number your CTO put in a quote.

Effective length

How many tokens it can actually hold accuracy across.

Much lower. Different per task. Never the same as the advertised number. This is the gap where teams get burned.

The point of the next thirty: three problems, three fixes, one budget. Five passes to trim an overloaded context, one budget split you can steal, and a rule that tells you when to compact, note, or spawn a sub-agent.
TAI Labs
By the end of this session

Five passes. One budget split. Three levers for when the budget blows.

The receipt: the five-pass trim (kill duplicates, keep rule, compress, reposition, trim tools), the starting budget split you can steal, and the decision rule for compaction vs. note-taking vs. sub-agents.
01
Name what makes context toxic. Attention is n². Lost-in-the-middle. Distractors. Agent exhaust.
02
Run the five-pass trim on an overloaded context. Duplicates. Keep rule. Compress. Reposition. Tools.
03
Budget the window on purpose. Instructions, task-critical facts, examples, working state, headroom.
04
Reach for the right lever when it blows. Compaction, note-taking, sub-agents.
TAI Labs
Why more context hurts · 01

Attention is n squared.

n tokens

n² pairs

Every token attends to every other token. That's your baseline attention budget.

2n tokens

4n² pairs

Double the context and you quadruple the pairwise relationships. Same fixed budget. Thinner attention on each pair. Not smarter. Just spread.

The physics of the thing: the transformer's attention cost scales quadratically with sequence length. The engineering budget behind that attention doesn't scale with your prompt. More context is not more focus. It's less.
TAI Labs
Why more context hurts · 02

Context rot and lost-in-the-middle.

Retrieval accuracy is high at the start and end of the window, and sags in the middle. Same fact, different position, different answer.

Accuracy Start Middle End lost here
Where you put the fact matters as much as whether it's there. The middle of a long window is where information goes to be forgotten. Design accordingly.
TAI Labs
Why more context hurts · 03

Distractor sensitivity. Near-miss is worse than nothing.

One correct chunk

Model has a single authoritative source.

Higher chance of the right answer. No ambiguity about which document to believe. Nothing competing with the truth.

Three plausible chunks

One is correct. Two look right.

The model can't tell which is authoritative. Accuracy drops. This is why "top-k retrieval" without conflict handling gets worse, not better, as k rises.

Stuffing "maybe relevant" material is not free. It actively competes with the truth. Every plausible-but-wrong chunk pulls the model's attention away from the one you meant it to use.
TAI Labs
Why more context hurts · 04

Needle-in-a-haystack lies to you.

What NIAH tests
  • Find a planted sentence buried in text.
  • Lexical overlap with the question is direct.
  • A retrieval trick, not a reasoning task.
What real work needs
  • Multi-hop reasoning across scattered pieces.
  • Conflicting material to reconcile.
  • Judgement under ambiguity, not lookup.
Passing NIAH at 500k tokens does not mean the model can reason over 500k tokens. Real work degrades much earlier than the NIAH curves suggest. Trust benchmarks that measure the shape of work you're actually doing.
TAI Labs
Why more context hurts · 05

Agentic loops fail differently.

Single-shot QA is one big blob. An agent accumulates its own exhaust.

Turn 1 task Turn 8 tools Turn 20 retries Turn 30 stale reads Turn 40 mostly exhaust
By turn 40, most of the context is a record of things that didn't work. The window is full of failed tool calls, stale reads, and retry cycles. The signal-to-noise ratio is fatal, and the model is trying to reason across all of it.
TAI Labs
Self-diagnose

Four signs you have a context problem, not a prompt problem.

01

Answers get vaguer as documents grow

Same prompt. Same question. Adding more retrieved material makes the response softer, less specific, less useful.

02

Wrong section, cited confidently

The model quotes a plausible but incorrect passage. It's not making things up. It's being misled by a distractor in the window.

03

An instruction that used to work stops working

You didn't change the prompt. You changed the surroundings. The instruction is drowning in nearby noise.

04

Cost climbs faster than usefulness

Each request keeps loading more into the window. Bills climb. Answers stay flat or degrade. That's a distraction tax.

Two or more of these are true? You don't have a prompt problem. You have a context problem. No amount of prompt polish will fix it.
The rule of the session

Find the smallest set of tokens that maximises the chance of the right answer.

Not the most context. Not the biggest window. The smallest high-signal set. Every token you add costs attention. Every token you remove has to actually be dead. This is the whole discipline in one sentence.

TAI Labs
The live trim

Cut an overloaded context in five passes.

01
Kill duplicates and stale versions. The most common cause of wrong-tool-call errors.
02
Apply the keep rule. If removing it wouldn't change the answer, it's decoration.
03
Compress, don't delete. Trade full content for summaries, schemas, diffs, signatures.
04
Reposition. Task and critical constraints at top and bottom. Middle is the amnesia zone.
05
Trim the tool surface. Twenty overlapping tools degrade selection accuracy on every turn.
Measure after each pass. Token count on one axis, output quality on the other. This is the whole shape of the session. Failure teaches the keep rule better than success does.
TAI Labs
Pass 01

Kill duplicates and stale versions.

Before

Three configs walk into a window.

config.v1.json config.final.json config.FINAL2.json

The model will pick the wrong one with confidence. It has no way to know which is current, and the file names actively lie to it.

After

One current source of truth.

Everything else is gone or archived outside the window. The model can only cite the version that's actually current.

Two versions of the same config file is the single most common cause of the model doing the wrong thing. Not model failure. Not prompt failure. Context failure. Cheapest possible fix, biggest possible payoff.
TAI Labs
Pass 02

Apply the keep rule.

The keep rule

If removing it wouldn't change the answer, it's decoration.

Say it out loud. Apply it to every block of context, one at a time. Decoration goes. The exam is: would the answer change?

What stays
  • Constraints the answer must satisfy.
  • Decisions already made, that this call inherits.
  • Facts the answer depends on.
  • Nothing else.
Emotional context is decoration. "For context, we've been talking about X for weeks." Model doesn't need that. Cut it. Every "just in case" is a bet against the keep rule, and the house wins.
TAI Labs
Halfway through

We stop. We hear the room. We keep going.

What we do
  • Drop one block of context in your agent that survives the keep rule against your better judgement.
  • One line. Just the type of content, not the whole document.
  • Three or four get read aloud.
  • Then straight into passes 3 through 5.
What this is for

Most of what you'll drop is transcript-shaped.

Full chat history, full tool output, full recall of prior conversations. All of it is decoration disguised as memory.

The next three passes give you the moves that keep the memory but drop the token weight.

TAI Labs
Pass 03

Compress, don't delete.

Instead of
Keep
Full transcript
Summary of decisions and open questions
Sample rows
Schema plus two example rows, max
Both file versions
The diff between them
Full implementation
Function signature plus docstring
Compression is not deletion. The information is still there. It's smaller. The model still has what it needs to reason. You just stopped paying attention tax on the boilerplate.
TAI Labs
Pass 04

Reposition. Top and bottom are high-attention zones.

Task + critical constraints Retrieved facts · working state · examples Repeat critical constraints
Never bury the instruction mid-window. Middle is where attention sags (see slide 5). Task at the top, restate at the bottom. If the model forgets what it's doing halfway through a long call, that's not the model. That's the layout.
TAI Labs
Pass 05

Trim the tool surface.

Twenty overlapping tools

Every definition is context tax on every turn.

You pay tokens for every tool description whether the agent uses it or not. And tool selection accuracy degrades as the menu grows. More options, worse choices.

Dynamic or collapsed tools

Load tools per step. Or expose a small allowlist.

Planning and execution should not share the same surface. Give the planner one set. Give each executor a scoped, minimal set. Selection accuracy jumps. Token spend drops.

The pattern: collapse tools by phase, or load them just-in-time based on the task class. Not every agent needs every tool on every turn.
TAI Labs
The live trim · past the edge

Cut too far. On purpose.

01
Remove one fact the answer actually depended on. Not decoration. Not "just in case." A real load-bearing fact.
02
Watch the model invent or miss. This is the failure. Feel it. Name it. It's what separates the trim from a blog post.
03
Say what you removed out loud. Then you know exactly what "load-bearing" means for this task. That is the keep rule made concrete.
The failure teaches the keep rule better than the success. Every trimmer overshoots eventually. Once. Then they never guess about what's decoration again. Better to overshoot here than in production.
TAI Labs
Budget the window

A starting split you can steal.

Instructions
10%
Task-critical facts
40%
Examples
15%
Working state / history
25%
Headroom
10%
When something comes in, something goes out. A budget with no eviction is not a budget. Task-critical facts get the biggest slice. Headroom protects you from the day a tool result comes back three times bigger than you expected.
TAI Labs
Budget the window · the retrieval move

Just-in-time retrieval. Over pre-loading.

Stuff everything up front

You guess what the agent will need.

The window fills with maybes. Half of them go unused. All of them steal attention from the ones the agent actually reaches for.

Let the agent go get it

Filesystem, grep, and targeted retrieval as context.

Navigation beats speculation. The agent asks for what it needs, when it needs it. The window stays tight. The reasoning stays fresh.

The mental shift: retrieval isn't a preprocessing step you run before the model. It's a tool the model uses. That's a different budget line and a different design.
TAI Labs
When the budget blows

Three levers. Each does something different.

01

Compaction

Summarise and reinitialise. Keep decisions and constraints. Discard redundant tool output. Fastest, cheapest, most disruptive to cache.

In-place
02

Note-taking

Write state to a file outside the window. Pull it back on demand across thousands of steps. Cache-friendly. Cheap on tokens; costs a filesystem.

External
03

Sub-agents

A specialist burns 10k tokens in a clean window and returns a 1.5k summary. Real gain when branches are independent. Coordination overhead when they aren't.

Parallel
The mental picture: a sub-agent gets a clean 10k window, does the work, returns a distilled 1.5k summary into the main agent's window. That's a 6-to-1 attention compression on that branch, at the cost of one extra model call.
TAI Labs
When the budget blows · pick the lever

Decision rule. Situation to move.

Situation
Reach for
Conversational flow that needs continuity
Compaction. Keep the thread coherent.
Long iterative work with milestones
Structured note-taking. The notes are the milestone log.
Parallel research threads
Sub-agents. One per thread, distilled back to the main.
The cache-cost caveat: aggressive trimming that reshuffles your prefix kills cache hits. The cheapest context is often the one you do not change. Weigh cache economics against attention economics on every call.
TAI Labs
What you do this week

Three things you can do by this weekend.

01
Log token counts per turn on one agent. You cannot cut what you cannot see. Simplest instrumentation. Highest leverage. Print input_tokens + output_tokens per turn to a file. That's the whole ask.
02
Delete your worst duplicate. The one you knew was there. Re-run. Measure the delta on token count and answer quality. Pass 1 of the trim, applied to real work.
03
Pick one raw tool dump and truncate or paginate it. The biggest tool result your agent returns. Compress it. Pass 3 of the trim, applied to the biggest offender.
Leave with one sentence: find the smallest set of high-signal tokens that maximises the chance of the right answer. Say it before every retrieval decision. Say it before every prompt edit. It rewires the whole discipline.
TAI Labs
The habit to take with you

This session was one call. Agents are where it gets hard.

The window refills every turn, and nobody is steering it unless you build the harness. Context engineering as a first-class skill, not a tip. Retrieval, memory, evals, and agent loops you can measure. Shipped systems, not slideware.

Next

Agentic AI Engineering Bootcamp

Nine weeks. Context engineering, evals, memory, and production agent loops. Certificate for engineers and AI PMs.

Where this leads
Now

Questions in chat

Post one line: the biggest context block you're afraid to trim. Aki or Manu will read a few aloud and diagnose live.

Open floor
Later

Your trim log

Instrument one agent. Log tokens per turn. Delete the worst duplicate. Send us the before/after. We reply with a short audio review.

The receipt
01 / 24