The return on AI training is measurable when three things exist: a baseline taken on real tasks before the first session, a like-for-like comparison of the same tasks after it, and assessed work rather than attendance as the evidence of capability. Most companies have none of the three, which is why a quarter of enterprise leaders say they cannot measure training ROI at all. This guide sets out the measurement method TAI Labs builds into every programme, the executive readiness report it produces, the per-team measures that make sense, and the numbers to be suspicious of.
Why the usual numbers do not work
Completion rates measure attendance. Seat activations measure curiosity. Self-reported time savings measure optimism. None of them survive a CFO's second question. The reason is that they measure the training, not the work. A return only exists if a task the company already does now takes less time, produces fewer errors, or reaches a standard it did not reach before, and that can only be seen by measuring the task itself, before and after, on comparable cases.
- Completion: says who watched, not who can do the work.
- Usage per person: creates gaming and resentment, and says nothing about output quality.
- Self-reported hours saved: inflated by enthusiasm, deflated by scepticism, never audited.
- Tool count: a team that learned one workflow well beats a team that opened twelve tools once.
The measurement method, step by step
Choose the workflow before you choose the measure
Measurement follows the work. Pick the one or two repeated workflows the training exists to change, and measure those. Company-wide productivity is not a measure; minutes to a reviewed account brief is.
Take the baseline on real tasks
In the week before the first session, time three to five real instances of each workflow and grade the output against a simple rubric (complete, accurate, sources present, ready to use). Record who did it and on what. This takes an afternoon and it is the step everyone skips.
Define a quality measure alongside the time measure
Faster and worse is not a return. For every time measure there is a quality measure: fields completed in the CRM update, revisions against the brand checklist, regressions per change, decisions linked to evidence.
Assess the work, not the attendance
Every participant produces assessed work during the programme: a workflow that runs, a skill file, a brief a reviewer accepted. Capability is evidenced by that work, scored against the rubric, not by a certificate of completion.
Compare like for like at week twelve
Repeat the baseline tasks on comparable cases, with the same rubric and, where possible, the same reviewer. Report the change per workflow with the sample size beside it.
Add an adoption measure that cannot be gamed
Not usage per person: the share of the workflow's real instances that went through the new process in the last two weeks, taken from the work itself (briefs produced, tickets created, exceptions logged).
Write the readiness report for the person who signs the next cheque
One page per team: baseline, week-twelve result, sample size, the assessed work behind it, and what to do next. Estimates labelled as estimates. The programme's executive readiness report is this document.
What to measure, by team
| Team | Time | Quality | Adoption |
|---|---|---|---|
| Sales | Minutes to prepare a customer meeting | Completeness of reviewed CRM updates | Briefs produced through the workflow, per week |
| Operations | Minutes to prepare a handover or report | Missing information caught before work starts | Exceptions logged and resolved |
| Marketing | Approved brief to reviewed concepts | Revision rounds against the brand checklist | Experiments launched and evaluated |
| Product | Research to a reviewed brief | Recommendations linked to evidence | Prototypes tested with real users |
| Engineering | Time to a reviewed, tested change | Review rework and regressions | Evaluation score for a shipped AI feature |
Turning hours into money, carefully
Leadership will ask for a figure in currency. The honest way to give one is capacity value: hours recovered per person per week, times people, times a loaded hourly rate, labelled as capacity rather than cash. A 20-person team recovering one hour each per week is roughly 1,000 hours a year of capacity. Whether that becomes revenue, headcount avoided or simply better work is a management decision, not a training outcome, and the report should say so. OpenAI's enterprise report found frontier companies route seven times more AI work through structured workflows than median ones; the capacity shows up where the workflows are, which is why the measurement is per workflow.
The tools for measuring it
- Google SheetsThe baseline log: task, minutes, rubric score, reviewer
- ClaudeScore assessed work against the rubric, with a person checking
- NotionThe rubric, the report template and the assessed work
- LinearEngineering time-to-change and rework, from the tickets
- HubSpotSales adoption: briefs and updates that went through the workflow
- AirtableThe operations exception log
Do it with us once
We set the baseline with you, then prove the change.
Every TAI Labs programme starts with a baseline and ends with an executive readiness report scored from assessed work. Bring one workflow to a 15-minute call and we will show you the measurement plan.
Numbers to be suspicious of
- Any percentage improvement quoted before the baseline exists.
- Vendor case studies with time saved but no sample size, no task and no reviewer.
- Company-wide productivity gains attributed to a training programme in one department.
- Adoption dashboards per employee presented as evidence of capability.
- Survey answers to 'how much time did AI save you this week'.
What the executive readiness report contains
- Per team: the workflows chosen, the baseline and the week-twelve comparison, with sample sizes
- The assessed work behind each score, with the rubric
- Adoption from the work itself, not from tool logs
- Capacity value, labelled as capacity, with the assumptions written down
- What to do next: the next workflow, the next team, the champions
- What did not work, and why
Frequently asked questions
How long before AI training shows a return?
A workflow measure moves within the programme, usually by the third live session, because the team is doing the real task the new way. A financial return depends on what management does with the capacity, and is best reported a quarter after the programme.
Can we measure ROI on training we already ran?
Partly. Without a baseline you can still take one now and compare a second cohort against it, or compare teams that were trained with teams that were not, on the same tasks. It is weaker evidence than a proper baseline, and the report should say so.
Should we track usage per employee?
No. It measures curiosity, invites gaming and reads as surveillance, especially to engineers. Measure the work: instances of the workflow that went through the new process.
What sample size is enough?
Enough to be honest about. Three to five real instances per workflow per person is typical for a baseline; report the number beside every figure and avoid percentages on samples too small to bear them.
Does TAI Labs guarantee a result?
No honest provider can. We guarantee the method: a baseline, assessed work, a like-for-like comparison and a readiness report with the numbers in it. The custom AI training page describes the programme.
Measure the task, not the training. Take a baseline on real work before the first session, pair every time measure with a quality measure, assess the work people produce, compare like for like at week twelve, and report capacity as capacity. That is how TAI Labs' executive readiness report is built, and it is why the programme is worth doing once, with the numbers to prove it.