
Aki Wijesundara
Manu Jayawardana
Tokens per conversation drift up for a week before a single ticket lands. Nobody watches this on a PM dashboard.
LeadingGroundedness dips on the top intent. Users still convert, still smile in CSAT. The regression is already happening.
LeadingBy the time churn shows up in your funnel, the feature has been broken for a month. Lagging metrics tell you when it is too late.
Lagging



Did the assistant finish what the user came for. Segmented by top five intents, not overall.
HealthPer intentShare of assistant claims that quote a real span from a real source. Drops here show up two weeks before hallucination tickets.
LeadingRAGWhere the assistant hands off. Watch by intent. A spike on one intent is a signal; overall averages hide it.
LeadingUXNot tokens. Not total spend. Cost per outcome. Rising cost per resolution is a quality problem wearing a finance jacket.
LeadingCost


Both are in your logs already. A Metabase card, a Looker tile, a single SQL query. If your team owns the assistant, you own these numbers today.
From logs you haveRule classifier on the top five intents. LLM judge on a daily sample of one to two hundred. Not perfect. Good enough to move a decision.
Sampled, LLM-judgedCite-span check against retrieved chunks. Now you can watch drift. Now the tile leads by a week, not lags by a week.
The one that leadsEvery tile on your dashboard needs a decision attached to it. If groundedness drops three points, what do you do. If it drops seven, what do you do. If you don't know, the tile is decorative. Take it off. Put a real one in its place.


Opt-in only. Users who thumbs-up already succeeded. Users who churned never voted. You are measuring your happiest cohort against itself.
Selection biasGoes up because usage went up. Goes down because usage went down. Tells you nothing about quality. Belongs on the finance dashboard, not this one.
ConfoundedLonger is not better. Longer means the user tried the assistant, gave up, tried again, gave up. Engagement and confusion look identical in this number.
AmbiguousThe refund intent is 3% of traffic and 40% of the pain. An "overall accuracy" tile averages it into oblivion. You will find the failure in the ops thread, not the dashboard.
Aggregated away
That is not a failure. It is the point. Lagging metrics feel safer because they are what leadership asks about.
The next block gives each of them a leading twin. Same feature. Different tile.

Five tiles, always visible:
Task success · Groundedness · Escalation · P95 retrieval latency · Cost per resolved chat.
The vitalsWhich intent changed shape. Which prompt got longer. Which source doc went stale.
Ranked by delta, not absolute value.
The storyOpen incidents. The current ship, hold, or roll-back call. The person who owns it.
A dashboard without an owner column is an inbox nobody reads.
The call

No P0 intent regressed. Cost per resolved chat within noise. Ship the feature to full traffic. Keep the tiles.
GreenGroundedness down three points on the refund intent. Root cause unclear. Stay at current traffic. Assign an owner. Re-check in two more days.
AmberYou will not diagnose faster than you can revert. Roll back the change. Then read the traces.
Red



Dashboards go stale. Traces do not. The PMs whose AI features age gracefully are the PMs who look at three real chats a week from their own product. Barely a coffee. Single highest-leverage thing on the calendar.
Nine weeks. Build the dashboard, the graders behind it, and the eval suite that guards them. Cohorts start monthly.
Where this leadsPost your feature's biggest quality risk in one line. Aki or Manu will read a few aloud and answer live.
Open floorDownload at tailabs.ai/dashboard-template.xlsx. Fill it. Email us the file. We send back a short audio review before the bootcamp starts.
The receipt