TAI Labs
TAI Lightning Lesson · Free

Deploy ML
from lab
to prod

The model is 10% of the system. The other 90% is the path it travels.
Session
Live · one sitting
You bring
A notebook · one model you want in production
Led by
Aki WijesundaraAki Wijesundara
Manu JayawardanaManu Jayawardana
TAI Labs
Why we are here

The gap between notebook and production is a systems gap.

Familiar

0.91 offline

Your best model, on your best split, in your notebook. Beautiful numbers. Slack screenshot ready. Impress the room.

Demo
The drop

0.72 a week later

Same model. Same data source. Different behaviour once it's serving real requests. Nothing modelling-related fixes this.

Production
Today

Fix the path

Three pieces of infrastructure close every gap that made the number drop. Not model tricks. System design.

The real move
The point of the next thirty: a six-question checklist you can run against your own project, three pillars that solve the three gaps, and one sentence you'll never forget.
TAI Labs
By the end of this session

Three pillars. Six checklist questions. And one sentence you can quote in every launch meeting.

The receipt: ML pipelines, feature stores, and CI/CD for ML · each mapped to the gap it closes. The six-question checklist scores you today. The bootcamp closes the ones you scored red.
01
Name the three gaps. Reproducibility, automation, consistency. Each needs its own fix.
02
Learn the three pillars. Pipelines, feature stores, CI/CD for ML. One per gap.
03
Watch the loop close. Monitor → trigger → retrain → shadow → promote. That's the ongoing system.
04
Run the six-question checklist. Six yes/no questions. Fewer than four yeses tells you where to start.
TAI Labs
The frame · three gaps

Three things stand between notebook and production.

Reproducibility

Same everything, every time

Same commit, same data hash, same artefact. Every time. Not "I think I trained this on the June split." Provable, not remembered.

Gap 1
Automation

No human zip-and-email

The path from data change to a serving artefact runs itself. Pipelines carry data, transforms, and models end to end.

Gap 2
Consistency

Train and serve compute the same features

If training saw a feature computed one way and serving computes it another way, the model has already left the lab. Silently.

Gap 3
You leave with: a six-question checklist you can run against your own project. Feature stores, pipelines, and CI/CD each solve one of these three gaps. Everything else in this deck is showing you how.
TAI Labs
The failure mode · four flavours

0.91 offline. 0.72 a week later.

01

Training / serving skew

The feature ran through pandas at train time. A different transform runs in the microservice at serve time. Same name, different meaning. The model quietly sees a different distribution.

Consistency
02

No lineage

Nobody can say which data produced which model. The commit hash exists somewhere. The data hash was never captured. Reproducing yesterday's model becomes an anthropological expedition.

Reproducibility
03

Manual handoff

Data scientist zips a pickle. Emails it to an engineer. Engineer wraps it in a service. Every step is a surface area for it worked in my notebook to become it doesn't work in prod.

Automation
04

No feedback loop

Nobody notices the drop for weeks. Then it's a P0. Then it's a re-training marathon. Because monitoring only saw the system up, not the system correct.

All three
The recurring line: "it worked in my notebook." None of these four are modelling problems. All four are systems problems. Three pieces of infrastructure fix all four.
The one rule to keep

"It worked in my notebook" is a systems failure. Fix the path, not the model.

Every fix has a target. When production diverges from the notebook, the model isn't the target. The path is. The three pillars are the three parts of that path. Everything else is variations on one of them.

TAI Labs
The system · three pillars

Three pillars. One for each gap.

1 · Pipelines

ML pipelines

Orchestrate ingest, validate, train, evaluate, register. Evaluation gates promotion. Every step is code, not a notebook cell run in someone's head.

Automation
2 · Features

Feature stores

One feature definition. Served offline for training and online for inference. Same transform, both places. Consistency by construction.

Consistency
3 · CI/CD for ML

Test code, data, and model together

Shadow, canary, then promote. With a rollback button. Because the change surface for ML is bigger than the change surface for software.

Reproducibility
Read across: pillar 1 automates the path, pillar 2 makes the path consistent, pillar 3 makes the path reproducible and reversible. Each pillar closes one gap. Skip a pillar and the corresponding gap reopens.
TAI Labs
Pillar 1 · Pipelines

From scripts to a DAG.

Ingest
raw data
→
Validate
schema + drift
→
Transform
features
→
Train
candidate
→
Evaluate
gate
→
Register
versioned
Eval gate

The red step is not automatic promotion

A model enters the registry as a candidate. Not as production. Metrics travel with it. The data hash travels with it. The commit travels with it. Later steps can trust or reject the candidate on evidence, not vibes.

Tools to know

Orchestration and registry, roughly in tiers.

Airflow Dagster Kubeflow Vertex SageMaker MLflow W&B

TAI Labs
Pillar 2 · Feature stores

One definition. Two serving paths.

A feature store is a single definition of a feature, materialised offline for training sets and online for low-latency lookups. Same transform, both directions.

Fixes skew

Write the transform once

Train and serve stop drifting apart by accident. The class of bug where a column called avg_spend_30d means one thing in the training set and another thing in the request handler simply cannot happen.

Reuse

Second team, no rewrite

The second team that needs "customer 30-day spend" reads the definition. Does not re-derive it. Every downstream model is now consistent with every other.

Honest caveat: one model with batch scoring probably does not need a feature store. Feature stores earn their keep with multiple models, multiple teams, or real-time serving. Not before.

Feast Tecton Cloud-native stores

TAI Labs
Pillar 2 · Point-in-time

Same transform. No future leakage.

Offline store

Historical features for training

Joined as-of the label timestamp. The value the transform would have produced at that moment, not today's value.

Training
Feature definition

One transform · versioned

The single source of truth for what this feature means. Referenced by both stores. Owned like code.

Shared
Online store

Latest values for inference

Low-latency lookups at request time. What the transform produces right now, materialised for fast reads.

Serving
Point-in-time correctness: if you join today's feature values onto last month's labels, you leak the future into training. Your offline metrics look great and then lie to you. Point-in-time joins are the fix.
TAI Labs
Halfway through

We stop. We hear the room. We keep going.

What we do
  • Drop your last "it worked in my notebook" moment in chat.
  • One line. The exact symptom. Not the fix.
  • Three or four get read aloud.
  • Then straight into pillar three and the loop.
What this is for

Most drops are one of the four

Skew, no lineage, manual handoff, no feedback. The taxonomy holds up. Naming the flavour narrows the fix.

Pillar three is the loop that catches those failures earlier next time. Then the six-question checklist scores where you are today.

The instinct to blame the model. Ignore it. The model did the best it could on the data it saw. The path is your surface area.
TAI Labs
Pillar 3 · CI/CD for ML

Test code, data, and model together.

Software CI/CD watches one input. ML CI/CD watches four. Any of them can silently invalidate the last model.

Code change Data change Schedule Drift signal

CI

Continuous integration

Lint, unit tests on transforms, data contracts, a fast smoke train on a sample. Cheap gate. Catches the silly stuff.

Fast
CT

Continuous training

Full retrain. Evaluate against the current production model, not a fixed threshold alone. Regressions are visible before they ship.

Grounded
CD

Continuous delivery

Register, shadow with no user impact, canary on a small slice of traffic, promote. Rollback is a button, not a scramble.

Reversible
TAI Labs
Pillar 3 · The loop

Monitor closes the circuit.

Monitor
inputs · preds · business
→
Trigger
drift · schedule · code
→
Retrain
CT pipeline
→
Shadow
compare silently
→
Promote
or roll back
Watch input drift first (fast), then prediction distribution, then the business metric (slowest, most important). Monitoring is the trigger back into the pipeline. Without monitoring, CI/CD/CT is a linear program, not a loop. Without the loop, quality decays and nobody sees it.
TAI Labs
Takeaway · the six-question checklist

Six questions. Score yourself honestly.

01
Can you reproduce your current production model from a commit hash? Reproducibility.
02
Is every feature defined in exactly one place? Consistency.
03
Does a failing data validation stop a training run? Automation with a real gate.
04
Can a new model reach production without a human copying a file? Automation.
05
Would you know quickly if input distributions shifted? Feedback loop.
06
Is rollback fast enough to be a real option? Reversibility.
Fewer than four yeses? Start with 1 and 3. Cheapest. Prevent the most damage. The rest earn their keep once those are in place.
TAI Labs
What you do this week

Run the checklist. Close one gap.

01
Score honestly. Six questions. Yes or no. No maybes. Whichever ones are no are the ones that will bite you first.
02
Pick the cheapest no. If reproducibility is broken, capture the data hash into the artefact registry. If validation isn't gating training, add one schema check. The pillar that fixes the cheapest no is where you start.
03
Write one short doc. Which gap, which fix, what will change on the next re-training run. Send it to whoever owns the pipeline with you. A gap you can name is a gap you can close.
04
Ship it before your next model iteration. The system fix compounds. The model fix doesn't. Fixing the path is the leverage move.
The receipt: the six-question score and the one paragraph. Send it. We open Week 1 of the Agentic AI Engineering Bootcamp by reading three of these together, choosing the highest-leverage fix per team.
TAI Labs
The sentence to remember

"It worked in my notebook" is a systems failure. Fix the path, not only the model.

Every deployment you own is a path with three sections: reproducibility, automation, consistency. Every drop from offline to online is a story about one of those three. Engineers whose models age gracefully are engineers who fix the path first and the model second.

Next

Agentic AI Engineering Bootcamp

Nine weeks. From models to production agents. Certificate for engineers and AI PMs. Cohorts start monthly.

Where this leads
Now

Questions in chat

Post one line: the "it worked in my notebook" moment you keep hitting. Aki or Manu will diagnose the gap live and name the pillar that closes it.

Open floor
Later

The score

Send your six-question checklist plus the one paragraph. We reply with a short audio review before the bootcamp starts.

The receipt
01 / 16