
Aki Wijesundara
Manu Jayawardana
Every check you wrote passes. Every screenshot lands. Every VC deck slide says green.
The first real user breaks it. Same model. Same prompt. Different outcome. The difference is the system around the call.







Does it parse? Are required fields present? Are types right?
Cheapest · always runValues in range? Referenced IDs exist? Does the arithmetic hold?
Cheap · run on every responseLLM or human check, only for what the first two cannot catch.
Slow · reach for last

Retries fail identically. You measure the same wrong answer three ways and think you have quality.
Wasted spendContext grows until quality degrades. Attention runs out before answers do.
Silent quality dropAn unbounded loop with write access. The blast radius grows on every iteration.
Money at riskYou faithfully persist a wrong answer and build on it. Later steps compound the mistake.
Compounding errorEvery "always" or "never" in your system prompt is a control you are hoping for instead of building. This is why failing prompts keep growing: each incident adds a sentence, the sentences compete, and past a point they degrade each other. The harness scales the other way. It is code, and code composes.





Not the prompt. Not the model. One of workflow control, context, permissions, evaluation, or state.
The next block is the side-by-side and the two tables that let you name the layer and the first fix.

Fits on a napkin. Ships in a demo. Breaks in production on the first unexpected input.
Schema validation · state · per-step context · tool allowlist · step budget · escalate path.
The model is identical. The prompt is nearly identical. One of these survives a real user, and the difference is roughly thirty lines that have almost nothing to do with AI.




Every failure your agent had this week maps to one of five layers. Before you touch the prompt, name the layer. Before you name the layer, run the diagnostic table. Engineers whose agents survive real users are engineers who stopped fixing at the wrong level.
Nine weeks. Build and ship production AI agents. Certificate for engineers and AI PMs. Cohorts start monthly.
Where this leadsPost one symptom. Aki or Manu will name the layer live and tell you the first thing to add.
Open floorTake the diagnostic table to your next incident review. Send us the write-up. We reply with an audio review before the bootcamp starts.
The receipt