TAI Labs
TAI Lightning Lesson · Free

The Agent
Harness

Why prompts fail in production. And what to build instead.
Session
Live · one sitting
You bring
One agent that broke in production
Led by
Aki WijesundaraAki Wijesundara
Manu JayawardanaManu Jayawardana
TAI Labs
The gap

The demo works. Production doesn't.

In testing

Same model. Same prompt.

Every check you wrote passes. Every screenshot lands. Every VC deck slide says green.

Week one in prod

Same model. Same prompt.

The first real user breaks it. Same model. Same prompt. Different outcome. The difference is the system around the call.

If you have shipped an agent that worked in testing and failed in production, we will map that failure to a specific layer of the harness by the end of the session.
TAI Labs
By the end of this session

Five layers named. Ten lines of harness. And a failure map that tells you what to build.

The receipt: the five-layer stack (workflow control, context, permissions, evaluation, state), the harness-or-prompt rule, a diagnostic table that maps every common failure to a specific layer, and the code shape that scales past the prompt.
01
Name the five layers. Every agent failure lives in one of them.
02
Learn the harness-or-prompt rule. Prompt says what good looks like. Harness says what happens when it isn't.
03
Read the code shape. Three lines to thirty. The extra twenty-seven have almost nothing to do with AI.
04
Map your own failure. Use the diagnostic table to name the layer and the first fix.
TAI Labs
Not the model. Not the prompt.

What actually changed between the demo and production.

Reality has
  • Inputs you did not imagine.
  • Users who change their mind mid-task.
  • Systems that time out.
  • Consequences when the agent is wrong.
A prompt cannot handle these
  • No memory.
  • No control flow.
  • No ability to stop itself.
  • A request, made once, with no enforcement.
The mental shift: a prompt is a request. A harness is a system. Your production agent is not one better prompt away from working. It's one better harness away.
TAI Labs
The harness

Five layers. The system around the model call.

01
Workflow control
What runs, in what order, and when it stops.
02
Context
What the model can see at this step.
03
Permissions
What it is allowed to touch.
04
Evaluation
How you know the output is good.
05
State
What survives between steps and between runs.
The claim: almost every agent failure lives in one of these five, not in the prompt. The next five slides prove it, one layer at a time.
TAI Labs
Layer 01 · Workflow control

When does this stop? Answer in code, not vibes.

01
A step budget, not vibes. Max iterations, max tool calls, max wall-clock.
02
A termination condition checkable by code. Not the model announcing "done."
03
A defined escape hatch. What happens on the final attempt when it still has not worked.
An agent that cannot stop itself is not autonomous, it is unsupervised. The common failure is not an infinite loop. It is the agent that declares success on step two because nothing was checking.
TAI Labs
Layer 02 · Context

What does the model see right now?

·
Context is assembled per step, not accumulated forever.
·
Belongs: current state, retrieved facts, the last relevant results.
·
Does not: the entire conversation history by default.
·
Stale context is worse than missing context. It looks authoritative, and the model cannot tell it is old.
Most teams treat context as an append-only log. Treat it as a query: "What does this step need to know?" every single time.
TAI Labs
Layer 03 · Permissions

What can it touch? Not "in principle." Right now.

·
Allowlist per step, not per agent.
·
Read and write are separate privileges. Most steps only need read.
·
Anything destructive, external, or irreversible gets a gate or a dry run.
·
The blast radius of a wrong decision is a design parameter, not an accident.
Classic failure: a tool kept write access because it was convenient in development and nobody removed it. Permissions failures are rare and expensive. Everything else is common and cheap.
TAI Labs
Layer 04 · Evaluation

How do I know this output is good? Three tiers, cheapest first.

01

Schema

Does it parse? Are required fields present? Are types right?

Cheapest · always run
02

Rules

Values in range? Referenced IDs exist? Does the arithmetic hold?

Cheap · run on every response
03

Judgement

LLM or human check, only for what the first two cannot catch.

Slow · reach for last
The dangerous failure is the plausible wrong answer: it passes casual review and reaches the user. Schema and rules are free and catch most of it. Judgement is slow, costly, and itself unreliable. Reach for it last.
TAI Labs
Layer 05 · State

What survives? What makes attempt two different from attempt one.

·
What has been attempted, and what failed, and why.
·
What the user has already confirmed or rejected.
·
Where we are in the workflow.
State is what makes attempt two different from attempt one. Without it, a retry is the same failing call again, at three times the cost. When an agent asks the same question twice, that is usually a state problem in your system, not a memory problem in the model.
TAI Labs
Five layers · how they interact

The layers are not independent. This is where most designs go wrong.

01

Evaluation without state

Retries fail identically. You measure the same wrong answer three ways and think you have quality.

Wasted spend
02

Context without workflow control

Context grows until quality degrades. Attention runs out before answers do.

Silent quality drop
03

Permissions without workflow control

An unbounded loop with write access. The blast radius grows on every iteration.

Money at risk
04

State without evaluation

You faithfully persist a wrong answer and build on it. Later steps compound the mistake.

Compounding error
Four out of five is not eighty percent of a harness. The missing layer is usually where the incident comes from.
The rule of the session

Prompt says what good output looks like. Harness says what happens when it isn't.

Every "always" or "never" in your system prompt is a control you are hoping for instead of building. This is why failing prompts keep growing: each incident adds a sentence, the sentences compete, and past a point they degrade each other. The harness scales the other way. It is code, and code composes.

TAI Labs
Code walkthrough · read, not run

Where everyone starts. Five failures hiding in three lines.

agent.py · the starting shape
def agent(user_input): result = model(user_input) return result
Malformed JSONEvaluation
TimeoutsWorkflow control
Invented fieldsEvaluation
Forgets two steps agoState
Forbidden tool callPermissions
Five failures. Five layers. That is not a coincidence. Every one of them is invisible in the code above, and every one of them will bite in production.
TAI Labs
Code · Evaluation

Move the guarantee into code.

agent.py · add a schema
from pydantic import BaseModel class Result(BaseModel): action: str target_id: str confidence: float raw = model(user_input) parsed = Result.model_validate_json(raw) # throws on malformed
You just moved "always return valid JSON" out of the prompt and into code that enforces it. The prompt asks. The schema guarantees. That distinction is the whole session.
TAI Labs
Code · State & context

Why attempt two can succeed.

agent.py · retry with state that learns
for attempt in range(MAX_ATTEMPTS): raw = model(build_context(state)) try: parsed = Result.model_validate_json(raw) state = state.update(parsed) break except ValidationError as e: state.log_failure(str(e)) # error goes back in else: return escalate(state)
The loop is not the interesting part. The interesting part is state.log_failure feeding into build_context: the only reason attempt two can succeed where attempt one failed.
TAI Labs
Code · Permissions & workflow

Allowlists are per step.

agent.py · scoped tool list + step budget
ALLOWED = {"search", "read_record"} # note: no writes result = model( context=build_context(state), tools=[t for t in TOOLS if t.name in ALLOWED], ) if state.steps > MAX_STEPS or state.is_complete(): return finalize(state)
The allowlist is a per-step decision. A planning step and an execution step should not have the same tools available. The step budget is what makes "return finalize" a real exit, not a hope.
TAI Labs
Halfway through

We stop. We hear the room. We keep going.

What we do
  • Drop one production surprise from your own agent, in chat.
  • One line. Not the diagnosis. The moment.
  • Three or four get read aloud.
  • Then straight into the diagnostic table.
What this is for

Most surprises map to one of five layers.

Not the prompt. Not the model. One of workflow control, context, permissions, evaluation, or state.

The next block is the side-by-side and the two tables that let you name the layer and the first fix.

The instinct to add another sentence to the prompt. Ignore it. Every "always" and "never" you write is a hint that a layer is missing. Add the layer.
TAI Labs
Code walkthrough · side by side

Three lines. Or thirty.

~3 lines
def agent(user_input): return model(user_input)

Fits on a napkin. Ships in a demo. Breaks in production on the first unexpected input.

~30 lines

Schema validation · state · per-step context · tool allowlist · step budget · escalate path.

The model is identical. The prompt is nearly identical. One of these survives a real user, and the difference is roughly thirty lines that have almost nothing to do with AI.

Order matters: schema first (nearly free), then rules, then anything that needs another model call. Cheap gates catch expensive failures. Reverse that order and you pay to detect problems you could have blocked.
TAI Labs
Apply the rule

Every "always" and "never" belongs somewhere else.

What you wrote in the prompt
Where it actually belongs
"Always return valid JSON"
Schema validation
"Never delete anything"
Tool allowlist
"Remember what the user said earlier"
State
"Do not make things up"
Retrieval + grounding check
"Stop when you have enough information"
Termination condition
"Be concise, use British spelling"
Stays in the prompt
Rule of thumb: if you are writing "always" or "never" in a system prompt, you are describing a control, not an instruction. Every "never" is a guarantee you are hoping for instead of building.
TAI Labs
Diagnose · map your failure

Six symptoms, six layers, six first fixes.

What you saw
Layer that broke
First thing to add
Agent looped forever
Workflow control
Step budget
Answered from stale data
Context
Rebuild context per step
Deleted the wrong record
Permissions
Per-step allowlist
Shipped a plausible wrong answer
Evaluation
Rules tier, not just schema
Forgot the user mid-task
State
Persist confirmed facts
Retried and failed identically
State + context
Feed the error back in
Print this. Stick it near your keyboard. Every incident postmortem starts by naming the layer. Every fix starts by adding the first thing on the right.
TAI Labs
Close · what we did not cover

Five things the harness also owns.

01
Evaluation sets and regression testing, so a change does not silently break step four.
02
Multi-agent handoff, where state becomes the hard part.
03
Human-in-the-loop gates and approval flows.
04
Observability: tracing a failure back to the step that caused it.
05
Cost and latency budgets as a harness concern, not a monitoring concern.
All of them extend the same five layers. None of them change the shape of the harness. The bootcamp covers each one with a running production agent.
TAI Labs
The habit to take with you

Prompt tuning has a ceiling. The harness does not.

Every failure your agent had this week maps to one of five layers. Before you touch the prompt, name the layer. Before you name the layer, run the diagnostic table. Engineers whose agents survive real users are engineers who stopped fixing at the wrong level.

Next

Agentic AI Engineering Bootcamp

Nine weeks. Build and ship production AI agents. Certificate for engineers and AI PMs. Cohorts start monthly.

Where this leads
Now

Questions in chat

Post one symptom. Aki or Manu will name the layer live and tell you the first thing to add.

Open floor
Later

The failure map

Take the diagnostic table to your next incident review. Send us the write-up. We reply with an audio review before the bootcamp starts.

The receipt
01 / 22