What is MLOps? The Machine Learning Lifecycle — lecture notes
Why a good model in a notebook is only the start: the ML lifecycle, hidden technical debt, training–serving skew, MLOps maturity levels, the tool landscape and the principles that keep models working in production.
0:001. Introduction

Training a model that scores well in a notebook is exciting, but it is only the beginning. A model must be deployed, monitored, retrained and trusted by real users for months or years. MLOps is the discipline that makes that possible. This deep dive gives you the big picture.
0:202. MLOps

MLOps is a set of practices and tools for reliably building, deploying, monitoring and improving machine learning systems in production. It extends the ideas of DevOps, which transformed software delivery, to the two things that make machine learning different: data and models.
0:383. Notebook vs production

The gap is large. In a notebook, the dataset is fixed and clean, the code runs once by hand, and success means a good test score. In production, data changes every day, everything must run automatically and continuously, and success means reliable value for users, month after month.
0:594. Hidden technical debt

A famous 2015 paper from Google, Hidden Technical Debt in Machine Learning Systems, drew this picture. The machine learning code is the small box in the middle. Around it sit configuration, data collection and verification, feature extraction, resource management, serving infrastructure and monitoring, and that is where most of the work lives.
1:215. The ML lifecycle

Machine learning work is a cycle, not a straight line. Frame the problem and define success. Collect, label and validate data. Train and evaluate against baselines. Deploy safely. Monitor data, predictions and outcomes. Then improve, which usually sends you back to the data. MLOps tries to make every turn of this cycle faster and safer.
1:446. DevOps vs MLOps

Why is DevOps not enough? In traditional software, behaviour is defined by code, and tests check that logic. In machine learning, behaviour is learned from data and code together, so we must also test the data and the model’s quality. And a model can silently decay even when nobody changes a line of code.
2:077. Changing anything changes everything

The same paper named the CACE principle: changing anything changes everything. Models entangle their inputs, so altering one feature, one data source or one threshold can shift behaviour everywhere. An upstream team renaming a product category can quietly degrade a model that nobody touched.
2:268. Pause and think

Pause and think. A model scored ninety five percent accuracy on its test set, but only about eighty percent after launch. What might have happened? Common causes are features computed differently in production, live data that differs from the training data, and leakage that made the original test score too optimistic.
2:479. Pause and think

Pause and think. You add a single new feature to a model. Following the changing anything changes everything principle, what should you re check? Everything: how existing features behave, performance on every segment, decision thresholds, latency, and any downstream systems that consume the predictions.
3:0610. Training–serving skew

One of the most common failures is training serving skew: features computed one way during training and another way at prediction time. Perhaps training used a thirty day average from a batch job, while the live service computes a seven day average. The model sees inputs it was never trained on, and quality drops.
3:2911. Who is involved

Many people are involved. A product owner defines the goal and how success is measured. Data engineers build reliable pipelines. Data scientists run experiments and produce models. Machine learning engineers turn them into production services, and platform engineers keep the infrastructure reliable. MLOps is as much about these handoffs as about tools.
3:5112. Maturity levels

Google Cloud describes three maturity levels. At level zero everything is manual: notebooks, hand run scripts and rare deployments. At level one, a training pipeline is automated so models retrain on new data. At level two, the pipelines themselves are built, tested and deployed automatically, just like application code.
4:1213. Start with a baseline

Before building anything complex, measure a baseline: a simple rule, such as predicting last week’s value, or a simple model like logistic regression. Ship the simplest thing that works, together with the pipeline around it, and iterate from there. Baselines often set a surprisingly high bar.
4:3114. Level 0

Level zero is where most teams start. A data scientist trains in a notebook, exports a model file and hands it to another team, who wire it into an application by hand. There is no monitoring, and retraining becomes a whole new project. It works for a first proof of concept, but it does not scale.
4:5515. Tracking experiments

A first step up is tracking every experiment. Each run records its parameters, metrics, code version and resulting model, so the best run can be found and reproduced later. Here five runs with different learning rates and dropout are compared, and the best one is clear.
5:1416. Automated pipelines

Higher maturity means automation. Every change triggers a pipeline that tests code and data, retrains the model, evaluates it against quality gates, registers it and deploys it gradually. The same steps run the same way every time, which makes releases boring, and in operations, boring is good.
5:3517. Versioning everything

Reproducibility requires versioning more than code. Data versions, code versions, experiment runs and model versions are linked, so for any model in production you can trace exactly which code, which data and which settings produced it, and rebuild it on demand.
5:5218. Packaging

Models must be packaged so they run identically everywhere. Containers bundle the operating system, libraries, model file and serving code into one image. The same image runs on a laptop, a cloud machine or a Kubernetes cluster, ending the classic excuse, it works on my machine.
6:1219. Serving

Once packaged, a model is served, often behind an API. Real traffic rises and falls through the day, so an autoscaler adds and removes replicas. When load spikes faster than capacity, latency rises above the service level objective, which is exactly what monitoring must catch.
6:3120. Monitoring

Monitoring watches more than servers. Here system metrics like latency and error rate look perfectly healthy, yet on day eighteen an upstream change breaks a feature. The distribution of predictions shifts immediately, and accuracy, measured days later when labels arrive, confirms the damage.
6:5021. Data drift

One reason monitoring matters is data drift. Here the live distribution of a feature slowly moves away from the training distribution over ten weeks. A drift score crosses a warning level and then an alert level, prompting investigation and retraining before accuracy quietly collapses.
7:0822. Feedback loops

Models can also influence their own future training data. A recommender that only shows popular items only collects feedback about popular items, so it grows ever more convinced they are the best. Such feedback loops can entrench bias, which is why teams keep some exploration and hold out traffic.
7:2923. Pause and think

Pause and think. A loan model rejects applicants it considers risky, so we never see whether they would have repaid. What does that do to retraining? The new training data contains only approved applicants, so the model can never learn it was wrong about the rejected groups. This selection bias needs deliberate exploration or careful correction.
7:5324. Principles

A few principles run through all of MLOps. Automate repeatable work with pipelines. Version code, data, models and configuration. Test the data and the model, not only the code. Monitor data, predictions and business outcomes. And make small, reversible changes, so that mistakes are cheap.
8:1225. Tool landscape

Many tools fill these roles. Git and DVC version code and data. MLflow and Weights and Biases track experiments and register models. Airflow, Kubeflow and Prefect orchestrate pipelines. FastAPI, BentoML, KServe and Triton serve models, and Prometheus, Grafana and Evidently monitor them. Principles matter more than particular tools.
8:3226. A tidy project

Good habits start with project structure. Data is versioned with DVC. A single feature module is shared by training and serving, preventing skew. Training, evaluation and serving are separate scripts that pipelines can run. Tests, configuration files, a Dockerfile and pinned dependencies make every run reproducible.
8:5227. Documentation

Documentation is part of the system. Model cards and datasheets for datasets describe what a model is for, what data it was trained on, how it performs, including for different groups of people, and what its limitations are. They make hand offs, audits and responsible use far easier.
9:1328. Cost awareness

Cost is a first class metric in production. Track the cost of training runs, GPU hours, storage and each prediction served, alongside accuracy. A one percent accuracy gain might not justify ten times the serving cost, and cheaper models can often be deployed more widely.
9:3229. Pause and think

Pause and think. Your team is at maturity level zero. What single step gives the most value first? Usually, make training reproducible: script it, version the data and track experiments. Then add basic monitoring. Full automation is much easier once these foundations exist.
9:5030. Common pitfalls

Watch out for common pitfalls. Some teams build an elaborate platform before shipping a single model. Others launch without a simple baseline to compare against. Many have no monitoring, so failures are silent. And models are often orphaned after launch, with nobody responsible for keeping them healthy.
10:1031. Measuring real impact

Finally, remember that offline metrics are only a proxy. The real test is impact on users, measured online, for example with an A B test comparing the current and new models on live traffic. Later in this track we cover how to run such tests correctly.
10:3032. Recap

To recap. MLOps extends DevOps to data and models. The model code is a small part of a real system. Beware training serving skew and feedback loops. Teams mature from manual work to automated pipelines to full CI and CD. And the core habits are to automate, version, test, monitor and iterate.
Key takeaways
- MLOps applies DevOps practices to data and models so ML systems keep working in production.
- In real systems, the ML code is a small part; data, serving and monitoring dominate (Sculley et al., 2015).
- Models can decay without code changes; changing anything can change everything (CACE).
- Training–serving skew and feedback loops are common, subtle failure modes.
- Maturity grows from manual (level 0) to automated training pipelines (level 1) to CI/CD of pipelines (level 2).
- Version everything, test data and models, monitor predictions and outcomes, and document with model cards.
Check yourself
- What makes MLOps different from traditional DevOps?
Show answer
Behaviour depends on data as well as code, and models can decay — Data and models add new failure modes.
- What is training–serving skew?
Show answer
Features computed differently in training and in production — The model sees inputs it was not trained on.
- At maturity level 1, what is automated?
Show answer
The training pipeline (continuous training) — Level 2 additionally automates CI/CD of the pipelines.
- A recommender that only shows popular items learns only about popular items. This is a…
Show answer
Feedback loop — Predictions shape future training data.
- Which is NOT a core MLOps principle?
Show answer
Deploy large untested changes at once — Prefer small, reversible changes.
Go deeper
- What Is MLOps? From Notebook to Reliable Production System · The AI Lecture Hall
- From Notebook to Production Code: Structuring Clean ML Projects · The AI Lecture Hall
- The Machine Learning Workflow: From Problem to Deployed Model · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/what-is-mlops-and-the-ml-lifecycle.html