TAI Labs
TAI Lightning Lesson · Free

Map the
4 layers
of AI eng

Find which layer your failure lives on. Then fix that layer.
Session
Live · one sitting
You bring
A stuck agent · one specific symptom
Led by
Aki WijesundaraAki Wijesundara
Manu JayawardanaManu Jayawardana
TAI Labs
Show of hands

Two weeks rewriting the system prompt. Bug still came back.

That was never a prompt bug. If more instructions have stopped helping, you are not on the prompt layer any more. The failure lives one layer up.
TAI Labs
Today

A map. A diagnostic. A clear layer.

By the end you can say out loud which layer your problem lives on.

1 · Prompt 2 · Context 3 · Loop 4 · Graph 5 · Questions
The move: a four-layer map, a climb rule you say out loud, a five-question diagnostic on your own agent, and a cheat sheet that maps symptoms to real layers.
TAI Labs
The line to remember

If more instructions have stopped helping, you are not on the prompt layer any more.

TAI Labs
The map

Four layers. Each solves what the one below cannot.

01Prompt
Wording, format, examples. One call.
invent facts it was never shown
02Context
Right info in the window, at the right time.
act, observe, then decide next
03Loop
Tools, retries, memory. Act, observe, revise.
cover true parallel branches alone
04Graph
Orchestration, workers, handoffs, checkpoints.
be worth ~15x tokens · multiplies failure modes
TAI Labs
How to climb

Name the gap. Then move one layer up.

If you cannot write one sentence for what the current layer cannot represent, stay put.

The sentence test: "Layer N cannot represent _____." If you can fill the blank, climb. If you cannot, the failure is still here and more work on this layer will fix it.
TAI Labs
Layer 1 · Prompt

Prompt work is real. Format alone moves the needle.

+76 pts
Same meaning, different format. LLaMA-2-13B (Sclar et al., 2023). GPT-3.5 spreads up to 56 pts on the same task.
Stay here when: the model already has the facts. Failure is wording, format, or order.
TAI Labs
Layer 1 · Ceiling

Rule 40 will not save you.

Still Layer 1

Everything needed is in the window.

Output shape is wrong. Tone slipped. A required field is missing. All fixable with wording, ordering, or a schema example.

Climb to Context

Confidently wrong about a fact it was never given.

No instruction fixes missing information. If you tell it more forcefully to be right, you get more forceful wrongness.

Sensitivity persists with bigger models and more few-shots. It is the interface, not a bug you instruct away.
TAI Labs
Layer 2 · Context

"Just use the big window." Not so fast.

99% → 70%
GPT-4o on NoLiMa: near-perfect in short context, ~70% at 32K. Ten of twelve frontier models halved their short-context score.
Chroma Context Rot: all eighteen frontier models degrade with length. Cliffs, not a gentle slope.
TAI Labs
Layer 2 · Engineering

Precision beats volume. Attention is the budget.

Still Layer 2

Answer exists in the corpus.

System misses it, or picks a lookalike. Quality drops as the thread grows.

Climb to Loop

Next step depends on an action that has not happened yet.

No pre-load can represent that. If the plan changes after seeing a tool result, no context strategy fixes it.

TAI Labs
Layer 3 · Loop

Act, observe, revise. Where most agents actually live.

Demo

Measures capability.

Can it do the task once?

Production

Measures reliability.

Can it do the task every time?

If you only report pass^1, you do not know what you built.
TAI Labs
Layer 3 · Ceiling

Capability is not reliability.

~25%

pass^8. GPT-4o agent, retail (tau-bench). ~60% drop from pass^1.

Reliability
80%+

Best models now clear pass^1. pass^k curves are still far from ideal (2026).

Capability
Still Layer 3

Works sometimes.

Runs diverge. Failures do not reproduce cleanly.

Climb to Graph

Work splits into independent branches.

One sequential trajectory cannot cover them.

TAI Labs
Layer 4 · Graph

Parallel structure pays when the work is parallel.

+90%
Anthropic Research: multi-agent beat single Opus 4 on research tasks. ~15x tokens. Breadth-first queries with independent directions.
Weaker on tightly interdependent work (e.g. coding). You are buying more context windows, not magic.
TAI Labs
Layer 4 · When it hurts

A graph multiplies failure modes.

41 · 87%
MAST: failure rates across seven multi-agent systems. Most failures are design (specification, misalignment, verification), not model size.
Pay only if branches are independent AND one correct answer is worth ~15x the tokens.
The rule of the session

If more instructions have stopped helping, you are not on the prompt layer any more.

Every fix has a layer. When your fixes stop working, that is not a signal to fix harder. It is the layer's ceiling, telling you to climb. Read the sentence again. Believe it.

TAI Labs
Halfway through

We stop. We hear the room. We keep going.

What we do
  • Drop one symptom from your own agent in chat.
  • One line. Not the diagnosis. The moment.
  • Three or four get read aloud.
  • Then we run the diagnostic on them, live.
What this is for

Most symptoms cluster on Layers 2 and 3.

The prompt-layer fixes and the graph-layer fixes are less common than they feel. The two middle layers are where the actual engineering lives.

The instinct to fix at the same layer that just failed you. Ignore it. When a layer's ceiling is real, more work there gets you nothing.
TAI Labs
Live · answer in chat

Five-question diagnostic

Q1
Expert writing by hand: what did the model lack?
Nothing → L1 Prompt · Doc → L2 Context
Q2
Does the next step depend on a previous action's result?
Yes → L3 Loop or higher
Q3
Same input, 8 runs. How many succeed?
Chain of steps collapses → L3 Loop
Q4
Do branches need each other's intermediate results?
No → L4 Graph maybe · Yes → stay L3 Loop
Q5
Is one correct answer worth ~15x your per-call spend?
No → stay L3 Loop
TAI Labs
Stop rule

One layer at a time.

Write the sentence for what this layer cannot represent. If you cannot, the problem is still here.

If more instructions have stopped helping, you are not on the prompt layer any more. Say it out loud. Believe it. Move up.
TAI Labs
Cheat sheet

Symptom → wrong fix → real layer.

You see
People fix
It lives on
Confidently wrong on facts
Prompt
Context
Quality dies in long chats
Prompt
Context
Demo works, prod flakes
Prompt
Loop
Ignores policy under pressure
Prompt
Loop
Too slow on broad research
Bigger window
Graph
Multi-agent worse than one
New model
Graph design
TAI Labs
Your turn

Drop your symptom in chat. One line.

We will tell you which layer it is on. This was the opening of week one. Bootcamp: build the loop with real evals. Measure pass^k on your system.

Next

Agentic AI Engineering Bootcamp

Nine weeks. Loop to production. Certificate for engineers and AI PMs. Cohorts start monthly.

Where this leads
Now

Questions in chat

Post one symptom. One line. Aki or Manu will diagnose the layer live and tell you where the fix lives.

Open floor
Later

The write-up

Send us the symptom, the layer, the sentence, and the fix. We reply with an audio review before the bootcamp starts.

The receipt
01 / 20