TAI Labs
TAI Lightning Lesson · Free · Build-along

A quality
dashboard
for PMs

One spreadsheet template. You leave with yours filled.
Session
Live · one sitting
You bring
A laptop · one AI feature you own · a Google account
Led by
Aki WijesundaraAki Wijesundara
Manu JayawardanaManu Jayawardana
TAI Labs
Why we are here

AI features degrade quietly. Your old dashboard will not tell you.

Cost

Climbs first

Tokens per conversation drift up for a week before a single ticket lands. Nobody watches this on a PM dashboard.

Leading
Quality

Wobbles second

Groundedness dips on the top intent. Users still convert, still smile in CSAT. The regression is already happening.

Leading
Users

Leave last

By the time churn shows up in your funnel, the feature has been broken for a month. Lagging metrics tell you when it is too late.

Lagging
The point of the next thirty: a sketch of the four leading numbers you can put on one page tonight. Enough to see the wobble before the users do.
TAI Labs
By the end of this session

A spreadsheet, filled with your feature. Ready to wire to real data this week.

The receipt: your own copy of the TAI Quality Dashboard template, four metrics chosen for the feature you own, thresholds written down, and one owner name in every incident row.
01
Pick metrics that predict breakage. Not vanity ones.
02
Separate signal from noise. Averages lie. Segments do not.
03
Design a dashboard PMs can act on. Every number has a decision behind it.
04
Ship, hold, or roll back. On the numbers. Not the vibes.
TAI Labs
Before anything else · grab your copy

One template. Save your copy. This is what we fill for the next thirty.

Do this now
01
Download tailabs.ai/dashboard-template.xlsx
02
Open it in Google Sheets, Excel, Numbers, whatever you have. Google Sheets: File → Import → Upload the file.
03
Save it with your feature name. Not "copy of template". A real name.
04
Post your feature name in chat. One line. So we know what you're working on.
What's in the Sheet
  • Tab 1 · Feature. Name, top five intents, baseline numbers.
  • Tab 2 · Health. Five KPI tiles. This is your top row.
  • Tab 3 · Drift. A weekly log of what changed shape.
  • Tab 4 · Decisions. Open incidents, owners, ship/hold/rollback.
  • Tab 5 · Thresholds. The green/amber/red bands per metric.
One template, pre-filled with the sample support assistant. As we go, you overwrite each cell with data from YOUR feature. By the close, the sample is gone. Yours is there.
TAI Labs
The feature we watch today

One support assistant. Seven days of production. Ten thousand chats.

Sample chat · day 4
User
what's your refund window?
Assistant
You can request a refund within 60 days of purchase. Just reply here and I'll start it.
User
i bought this 45 days ago, so i'm good?
Assistant
Yes, you're inside the window. Confirming your refund now.
What quietly happened
Policy is 30 days. The assistant invented a 60-day window three weeks ago.
CSAT stayed at 4.6. Users love the answer. Ops eats the refund.
Cost per chat rose 18%. Longer retrieval calls, more tool loops.
This is the feature. Every metric today lives around it.
TAI Labs
Block 1 · Pick your four

Four numbers that move before your users do.

01

Task success, by intent

Did the assistant finish what the user came for. Segmented by top five intents, not overall.

HealthPer intent
02

Groundedness on retrieval

Share of assistant claims that quote a real span from a real source. Drops here show up two weeks before hallucination tickets.

LeadingRAG
03

Escalation rate to human

Where the assistant hands off. Watch by intent. A spike on one intent is a signal; overall averages hide it.

LeadingUX
04

Cost per resolved conversation

Not tokens. Not total spend. Cost per outcome. Rising cost per resolution is a quality problem wearing a finance jacket.

LeadingCost
In your Sheet, tab 2 · Health, column A: write your four in the "metric" column. Not mine. Yours. If your feature has no groundedness signal because there's no retrieval, swap it for something that predicts breakage. Add a fifth for latency if speed matters to your users. The template Sheet includes p95 latency as an example. Drop your list in chat when done.
TAI Labs
The four, up close

What you actually measure. And where each one breaks first.

Metric
What you actually measure
Where it breaks first
Task success, per intent
Wanted-outcome rate on the top five intents. Rule classifier plus LLM judge on a daily sample.
New intents. Ambiguous queries. Anywhere the router is guessing.
Groundedness
Share of assistant claims that quote a real span from a retrieved source.
Policy doc updates. Stale retrievers. The day a source chunk goes out of date.
Escalation to human
Hand-off rate per intent. Watched on the top five and every new intent this month.
The intent that ships refunds. A prompt template that just went live.
Cost per resolved chat
Total spend divided by successful resolutions. Not per token. Per outcome.
Tool-loop regressions. A verbose model swap. Longer retrieval calls nobody noticed.
In your Sheet, tab 2 · Health, column B ("what you measure"): write the one-sentence definition for each of your metrics. If you can't write it in one sentence, the tile isn't ready to ship. Vague measurement is what makes dashboards lie.
TAI Labs
One chain, five links

Watch the left. The right will already be true.

RETRIEVAL QUALITY chunks match intent, policy doc is fresh watch this first GROUNDEDNESS claims trace back to a real source leads by ~7 days TASK SUCCESS user got what they came for, by intent leads by ~4 days ESCALATION RATE users give up on the assistant leads by ~2 days CSAT · CHURN · TICKETS the users finally tell you too late to prevent
The dashboard job: put the leftmost three on the top row. Anyone can watch the rightmost one. You need the ones that move earlier.
TAI Labs
How to actually get these numbers

You don't need a platform. You need a warehouse and a v0.

Ship 1

Cost and escalation

Both are in your logs already. A Metabase card, a Looker tile, a single SQL query. If your team owns the assistant, you own these numbers today.

From logs you have
Ship 2

Task success per intent

Rule classifier on the top five intents. LLM judge on a daily sample of one to two hundred. Not perfect. Good enough to move a decision.

Sampled, LLM-judged
Ship 3

Groundedness

Cite-span check against retrieved chunks. Now you can watch drift. Now the tile leads by a week, not lags by a week.

The one that leads
In your Sheet, tab 2 · Health, column H ("ship order"): mark each of your tiles 1, 2, or 3. Cheap first, hard last. Every ship gives you a real tile to argue with. A v0 in production beats a v1 in a doc.
The rule for the next thirty

If a number can't move you, it isn't on the dashboard. It's furniture.

Every tile on your dashboard needs a decision attached to it. If groundedness drops three points, what do you do. If it drops seven, what do you do. If you don't know, the tile is decorative. Take it off. Put a real one in its place.

TAI Labs
Block 2 · Signal vs. noise

Averages hide the failure. Segments show it.

Signal
  • +Task successtop 5 intents
  • +Groundednessrefund intent
  • +Cost per chatby tool-loop count
  • +Escalationnew intents this month
  • +p95 latencyretrieval step only
Noise
  • −CSATaveraged across intents
  • −Token spendtotal, not per outcome
  • −Session lengthall users combined
  • −Thumbs-upopt-in only
  • −Latencyend-to-end mean
The trick: the assistant that invented a 60-day refund window kept overall CSAT steady. It only showed up in groundedness on the refund intent. If your dashboard averages across intents, you would still be shipping it.
TAI Labs
The vanity metric graveyard

Four tiles that look responsible. And what they are hiding.

01

Thumbs-up rate

Opt-in only. Users who thumbs-up already succeeded. Users who churned never voted. You are measuring your happiest cohort against itself.

Selection bias
02

Total token spend

Goes up because usage went up. Goes down because usage went down. Tells you nothing about quality. Belongs on the finance dashboard, not this one.

Confounded
03

Average session length

Longer is not better. Longer means the user tried the assistant, gave up, tried again, gave up. Engagement and confusion look identical in this number.

Ambiguous
04

Overall accuracy

The refund intent is 3% of traffic and 40% of the pain. An "overall accuracy" tile averages it into oblivion. You will find the failure in the ops thread, not the dashboard.

Aggregated away
In your Sheet, tab 2 · Health: if any of your four rows is on this graveyard slide, delete it now. Replace it with a leading indicator. Real accountability is a number tied to a decision. If you can't finish the sentence "if this moves, I will…" the tile is vanity.
TAI Labs
Halfway through

We stop. We hear the room. We keep going.

What we do
  • Look at tab 2 of your Sheet. Read the four metric names you wrote.
  • Post the one that would move first for your feature, in chat.
  • Three or four get read aloud.
  • Then straight into the three-tier layout.
What this is for

Most of what you dropped is lagging

That is not a failure. It is the point. Lagging metrics feel safer because they are what leadership asks about.

The next block gives each of them a leading twin. Same feature. Different tile.

The instinct to defend the metric. Ignore it. The old number stays. We add the ones that move first, next to it.
TAI Labs
Block 3 · Draw the skeleton

One page. Three tiers. The whole dashboard.

Tier 1

HEALTH · today

Five tiles, always visible:

Task success · Groundedness · Escalation · P95 retrieval latency · Cost per resolved chat.

The vitals
Tier 2

DRIFT · this week

Which intent changed shape. Which prompt got longer. Which source doc went stale.

Ranked by delta, not absolute value.

The story
Tier 3

DECISIONS · now

Open incidents. The current ship, hold, or roll-back call. The person who owns it.

A dashboard without an owner column is an inbox nobody reads.

The call
Your Sheet has three tabs that match these three tiers. Health, Drift, Decisions. The rule of the skeleton: Tier 1 fits on a laptop, Tier 2 is a scroll, Tier 3 is a link. If Tier 1 scrolls, you are looking at the wrong numbers.
TAI Labs
The sample feature · pre-filled in the template

The whole dashboard fits in one spreadsheet.

refund-window-QA · quality dashboard quality-dashboard.xlsx · sample view FILE EDIT VIEW INSERT FORMAT DATA TOOLS EXTENSIONS HELP A · METRIC B · WHAT YOU MEASURE C · CURRENT D · LAST WK E · DELTA F · BAND G · IF-THIS-MOVES-I-WILL Task success · top 5 intents Outcome rate, top 5 intents 92.4% 92.1% +0.3 GREEN below 85% on any intent → open a HOLD Groundedness · refund intent Claims traced to a retrieved span 87.1% 90.3% −3.2 AMBER assign an owner, re-check in two days Escalation rate · per intent Hand-off rate per intent 7.8% 7.6% +0.2 GREEN any intent above 15% → open a HOLD Cost per resolved chat Spend / successful resolutions $0.14 $0.11 +18% AMBER check tool-loop count on the refund intent P95 retrieval latency P95 of retrieval only, not end-to-end 2.4s 2.5s −0.1 GREEN above 4s → open a HOLD on the retriever OPEN DECISIONS · see tab 4 · row 3 INC-142 · refund intent · groundedness + cost both amber Owner: PM-refunds HOLD · re-check day 4 1 · FEATURE 2 · HEALTH ● 3 · DRIFT 4 · DECISIONS 5 · THRESHOLDS
The whole thing is a spreadsheet. Not a dashboard tool. Not a platform. Eight columns, five rows, one decision per row. If you can fill this out for your feature, you have a dashboard. Everything after this is wiring it to real data.
TAI Labs
Three decisions the dashboard supports

Ship. Hold. Roll back. On the numbers.

Ship

All your leading tiles steady over seven days.

No P0 intent regressed. Cost per resolved chat within noise. Ship the feature to full traffic. Keep the tiles.

Green
Hold

One leading indicator crossed threshold.

Groundedness down three points on the refund intent. Root cause unclear. Stay at current traffic. Assign an owner. Re-check in two more days.

Amber
Roll back

Two indicators crossed together, or a P0 intent lost ten points.

You will not diagnose faster than you can revert. Roll back the change. Then read the traces.

Red
In your Sheet, tab 4 · Decisions, row 1 (the SESSION RULE row): overwrite the template rule with YOUR feature's version. "Ship if…, Hold if…, Roll back if…" Pre-commit. In the incident, nobody argues about whether three points counts. The Sheet already said.
TAI Labs
The threshold cheat sheet

Write the numbers before the incident. Not during.

Metric
Green
Amber · Red
Task success, top 5 intents
≥ 90% steady over the last week
Amber: any intent between 85 and 90%. Red: any intent below 85%.
Groundedness
≥ 92% steady
Amber: 88 to 92%, or dropping 3 points week over week. Red: below 88%, or dropping 5 points week over week.
Escalation, per intent
≤ 10% for every top-5 intent
Amber: any intent between 10 and 15%, or +3 pts week over week. Red: above 15%, or +5 pts week over week.
Cost per resolved chat
Within ±10% of your baseline
Amber: 10 to 20% above baseline. Red: above 20% above baseline.
P95 retrieval latency (optional 5th)
Within ±10% of baseline
Amber: 10 to 30% above baseline. Red: above 30% above baseline.
In your Sheet, tab 5 · Thresholds: paste these as your starting rows, then tune the numbers to your product. The 5th row (latency) is there as an example. Keep it if UX depends on speed. Drop it otherwise.
TAI Labs
Applying it · the refund incident, replayed

Same chat. A dashboard that would have caught it.

What actually happened
  • Day 0. Policy source doc updated: 30 days.
  • Day 1. Retrieval still hitting the old chunk. Groundedness on refund intent drops from 94% to 87%.
  • Day 3. Task success on refund intent holds. CSAT holds. Cost per resolved chat up 12%.
  • Day 7. Ops notices the refund count. The finance team notices before the PM.
What the dashboard would have said
  • Day 1. Tier 1 groundedness tile: amber.
  • Day 1. Tier 2 drift tile: "refund intent · retrieval source changed."
  • Day 2. Tier 3 decision: HOLD. Assign to the on-call PM.
  • Day 3. Root cause found in the traces. Roll back the retrieval config. Cost, quality, ops all recover.
In your template, tab 4 · Decisions, row 3. That's INC-142 as the sample sheet saw it in real time. Nothing new was measured. The dashboard is the difference between "the metric was collected" and "the metric was seen in time to act." That's the whole job.
TAI Labs
What you do this week

Wire your Sheet to one real data source.

01
Pick the cheapest metric in your Sheet. The one you marked Ship 1 in column H. Usually cost per resolved chat or escalation rate. You want the smallest possible round-trip from raw log to filled cell.
02
Get the raw data out. A warehouse export, a Metabase CSV, an =IMPORTDATA on a webhook, a Zapier from Zendesk. Whatever your stack allows in a working session. If you can't get it out yourself, book fifteen with your data engineer this week.
03
Fill one row of the Health tab with a real number. Not a mock. Not a placeholder. One real cell. That's the whole homework. The rest of the tiles follow the same pattern next week.
04
Share the Sheet with one engineer and one PM peer. Ask them: does this tile mean what it says. What's missing. A dashboard nobody else has seen isn't a dashboard.
The receipt: your filled spreadsheet with your feature's name, four metric rows, at least one wired to real data, and thresholds in tab 5. Email us the file (or the Sheets link if you moved it to Google Sheets). We'll open Week 1 of the AI Evals Bootcamp by looking at yours together.
TAI Labs
The habit to take with you

Read three traces a week. Not the failing ones. Random.

Dashboards go stale. Traces do not. The PMs whose AI features age gracefully are the PMs who look at three real chats a week from their own product. Barely a coffee. Single highest-leverage thing on the calendar.

Next

AI Evals Bootcamp

Nine weeks. Build the dashboard, the graders behind it, and the eval suite that guards them. Cohorts start monthly.

Where this leads
Now

Questions in chat

Post your feature's biggest quality risk in one line. Aki or Manu will read a few aloud and answer live.

Open floor
Later

The template

Download at tailabs.ai/dashboard-template.xlsx. Fill it. Email us the file. We send back a short audio review before the bootcamp starts.

The receipt
01 / 20