CI/CD for ML and Safe Deployment Strategies
Automate the path to production: testing code, data and models; CI/CD/CT pipelines with quality gates; champion–challenger evaluation; shadow, canary and blue–green deployments; and fast rollback.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
5 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. How do you change a model that millions of people rely on without breaking anything? You automate the path to production and release changes gradually, with the ability to undo them instantly. This deep dive covers continuous integration, delivery and training for machine learning, and the deployment strategies that keep releases safe.
CI, CD and CT. Three ideas work together. Continuous integration tests every change automatically. Continuous delivery, or deployment, releases changes automatically and safely. Machine learning adds continuous training, which retrains models automatically when new data arrives, because a model can go stale even when the code does not change.
A CI/CD pipeline. Here is a pipeline in action. A commit triggers unit tests and a data schema check. The model is trained on the latest data, then evaluated: accuracy must reach at least point nine and there must be no fairness regression. The model is registered, and finally deployed as a canary to five percent of traffic.
Tests for ML. Machine learning needs several kinds of tests. Unit tests check code logic. Data tests check schemas, ranges and freshness. Model quality tests compare metrics with thresholds, overall and for important segments. Behavioural tests check expected behaviours, and integration tests check the whole service end to end.
Behavioural testing. Behavioural testing, popularised by the CheckList method, treats a model like software with specifications. Invariance tests check that irrelevant changes, such as swapping a person’s name, do not change the output. Directional tests check that some changes move the output the expected way, like adding the word not.
Pause and think. Pause and think. Can you write an invariance test for a loan approval model? For example: change only the applicant’s name, or another attribute that should be irrelevant, while keeping every financial detail identical. The decision must not change. Failures reveal bias or brittle behaviour.
A failing quality gate. Now watch a quality gate do its job. Tests pass and training completes, but the new model scores only point eight seven, below the required point nine. The pipeline stops, the model is never registered or deployed, and the team is notified. Bad models are caught automatically, before users ever see them.
Quality gates. A quality gate is an automated checkpoint the pipeline must pass. Typical gates require minimum metrics, no regression compared with the model in production, fairness checks across groups, and latency within budget. Gates turn quality standards from good intentions into enforced rules.
A CI workflow. Here is a small CI workflow for GitHub Actions. On every push it checks out the code, sets up a pinned Python environment and installs dependencies. It runs the unit and data tests, trains the model with a versioned configuration, and runs an evaluation script that fails the build if accuracy is below the gate.
Continuous training. Continuous training needs triggers. Some teams retrain on a schedule, such as weekly. Others retrain when enough new labelled data has accumulated, or when monitoring detects drift or falling performance. Either way, an automatically retrained model must pass through exactly the same tests and gates.
A CT pipeline. A continuous training pipeline starts from a trigger, validates the new data, trains a challenger model and compares it with the current champion on the same evaluation data. If the challenger is better, it is registered for deployment. If not, the pipeline stops and alerts the team, and the champion stays in place.
Champion vs challenger. This champion challenger pattern keeps quality from sliding backwards. The model in production, the champion, is only replaced when a challenger beats it on agreed metrics: first offline on held out data, and then online, on a small share of real traffic.
Pause and think. Pause and think. The challenger is half a percent better offline. Should you switch all traffic to it immediately? No. Offline gains do not always hold on live traffic, and new failure modes can appear. Roll it out gradually while monitoring, with a rollback ready.
Shadow deployment. The safest first step is shadow deployment. A copy of live requests goes to the new model, but users only ever see the old model’s answers. You can compare predictions, latency and errors on real traffic with zero risk to users, catching crashes and training serving skew early.
Pause and think. Pause and think. A shadow deployment shows the new model’s predictions look sensible and fast. What can it not tell you? How users and business outcomes respond, because nobody ever acted on its outputs. You still need a canary release or an A B test to learn that.
Canary release. Next comes a canary release. The new version receives one percent of traffic, then five, then twenty five, fifty and finally one hundred percent. At each step, error rates and business metrics are checked against the old version. A bad release reaches only a few users before it is caught.
Automatic rollback. Here the canary goes wrong. At twenty five percent of traffic, the new version’s error rate climbs past the two percent threshold. The rollout is stopped automatically and all traffic returns to the old version. Only a quarter of users saw the bad version, and only briefly.
Blue–green deployment. Blue green deployment runs two complete environments. Blue serves users while green, with the new version, is tested. Then all traffic switches to green at once. If anything goes wrong, traffic switches straight back to blue. It gives instant rollback, at the cost of running double capacity during the switch.
Canary vs blue–green. Canary releases are gradual and expose only a few users at first, but they take longer and depend on good monitoring. Blue green switches everyone at once with instant rollback, but every user is exposed immediately. Many teams shadow first, then canary, keeping the previous version warm for a blue green style rollback.
Feature flags. Feature flags add fine control without redeploying. A configuration switch can turn a new model on or off, or route specific groups, such as employees or one country, to it. Flags make it easy to test with friendly users first, and to turn a problem off in seconds.
Online evaluation. To decide whether the new model is truly better, many teams run an A B test during the rollout, comparing a business metric such as conversion between the two versions. Here the new model’s advantage becomes statistically clear by the planned end of the test, and it is rolled out.
Build once, deploy many. A key rule is build once, deploy many. Build and test one immutable image, then promote that exact image through development, staging and production, changing only configuration. What you tested is exactly what you ship, so there are no surprises caused by rebuilding.
Rollback readiness. Always be ready to roll back. Keep the previous model version deployed or instantly deployable, together with the feature pipeline and configuration it expects. Practise rollbacks regularly. The goal is to be back on the old model within minutes, not days, whenever something looks wrong.
Slice-based validation. Validation must go beyond overall numbers. Compare the challenger with the champion on important slices: regions, device types, customer groups and rare but critical classes. A model can win on average while losing badly on a small but important segment, and slice checks catch that before users do.
GitOps. GitOps takes this further. The desired state of every environment, including which model version is deployed, lives in a Git repository. An automated agent such as Argo CD or Flux continuously makes the cluster match it. Deploying means merging a reviewed pull request, and rolling back means reverting it.
Infrastructure as code. Infrastructure itself should be code. Servers, clusters, pipelines and permissions are described in versioned files, such as Terraform configurations and Kubernetes manifests, instead of manual clicks in a console. Environments can then be reviewed, reproduced and rolled back just like application code.
Audit trail. All this automation leaves an audit trail. For every model in production, you can see which code, data and parameters produced it, which tests and gates it passed, who approved it and when it was deployed. In regulated industries such as finance and healthcare, this traceability is often a legal requirement.
Strategies summary. In summary: shadow deployments carry no user risk and catch crashes and skew. Canaries expose a small, growing share with automatic rollback and suit most model updates. Blue green switches everyone at once with instant rollback, and A B tests measure real business impact.
Pitfalls. Watch for pitfalls. Flaky tests, caused by unfixed seeds or tiny test sets, erode trust in the pipeline. Gates that are too loose, or on the wrong metric, let bad models through. Averages can hide regressions for particular groups. And big bang releases make it impossible to tell which change caused a problem.
Pause and think. Pause and think. Your model quality test fails randomly about one run in ten, even with no changes. What should you do? Fix the flakiness: set random seeds, use a larger fixed evaluation set, and choose thresholds that allow for measurement noise. Never simply re run until it passes.
Recap. To recap. Continuous integration tests code, data and models, and continuous training retrains them automatically. Quality gates block weak models. Champions are replaced only by clearly better challengers. Release gradually, from shadow to canary to full rollout, build once and deploy many, and always keep rollback ready.