Week 4: Evals, the TRACE loop (shared with Engineering cohort)
8 lessons · Back to full syllabus
What you keep
How to evaluate an AI feature like a leader, using TRACE. Error analysis is product work, you own the front of this loop. You can brief eng for Codify/Enforce with TPR/TNR, not vibe metrics.
You ship
Your product, evaluated, with a must-pass checklist and at least one fix before the sprint.
Live session resources
Shared TRACE session deck (Engineering Week 4)
Same masterclass deck as Engineering - golden sets, checks, judges, and shipping discipline.
Open resourceWeek 4 assignment support guide
Path A / Path B: TRACE report, must-pass checklist, one fix; stretch Braintrust annotations and engineer brief.
Open resourceSample traces pack (Harmony Apartments)
20 SMS leasing-bot traces + knowledge base (CSV for Sheets) when you do not have product traffic yet.
Open resourceHow to Build AI Evals in 2026 (YouTube)
Step-by-step evals framing for the vibe-check trap.
Open resourceMastering AI Evaluation: Playground to Production (YouTube)
From ad-hoc checks to production eval loops.
Open resourceBraintrust
Primary classroom tooling - annotate traces and add scorers.
Open resourceLangfuse
Open-source backup for tracing and evals.
Open resourceAnthropic: Demystifying evals for AI agents
Task, trial, grader, transcript vs outcome - language for briefing eng.
Open resourceLessons
The vibe-check trap
"It looked good when I tried it" is not evaluation, and vibes do not survive change.
TRACE: Trace and Read (error analysis is your job)
Capture real interactions, read them one by one, and journal what went wrong - product work, not engineering.
TRACE: Analyze (decide what matters, fix the obvious)
Cluster failures by frequency, fix the cheap ones, and judge pass/fail, not vague scores.
What good evals tooling looks like (so you can lead it)
Codify and Enforce are engineering, but you can recognise good checks and hold a team to them.
Build a simple must-pass checklist for your product
Turn top failures into binary pass/fail cases, and re-run the list every time you change the product.
How to brief an engineer to build the evals you need
Hand traces and must-pass cases, not a request for generic quality metrics.
Week 4 assignment support: Run TRACE on your product
Path A: traces, ranked failures, must-pass checklist, one fix. Path B: Braintrust annotations and engineer brief.
Run TRACE on your product
Traces read, failures ranked, must-pass checklist built, at least one fix shipped. Path B optional.
Lessons in this module