CI/CD for ML and Safe Deployment Strategies — lecture notes
Automate the path to production: testing code, data and models; CI/CD/CT pipelines with quality gates; champion–challenger evaluation; shadow, canary and blue–green deployments; and fast rollback.
0:001. Introduction

How do you change a model that millions of people rely on without breaking anything? You automate the path to production and release changes gradually, with the ability to undo them instantly. This deep dive covers continuous integration, delivery and training for machine learning, and the deployment strategies that keep releases safe.
0:212. CI, CD and CT

Three ideas work together. Continuous integration tests every change automatically. Continuous delivery, or deployment, releases changes automatically and safely. Machine learning adds continuous training, which retrains models automatically when new data arrives, because a model can go stale even when the code does not change.
0:413. A CI/CD pipeline

Here is a pipeline in action. A commit triggers unit tests and a data schema check. The model is trained on the latest data, then evaluated: accuracy must reach at least point nine and there must be no fairness regression. The model is registered, and finally deployed as a canary to five percent of traffic.
1:044. Tests for ML

Machine learning needs several kinds of tests. Unit tests check code logic. Data tests check schemas, ranges and freshness. Model quality tests compare metrics with thresholds, overall and for important segments. Behavioural tests check expected behaviours, and integration tests check the whole service end to end.
1:245. Behavioural testing

Behavioural testing, popularised by the CheckList method, treats a model like software with specifications. Invariance tests check that irrelevant changes, such as swapping a person’s name, do not change the output. Directional tests check that some changes move the output the expected way, like adding the word not.
1:446. Pause and think

Pause and think. Can you write an invariance test for a loan approval model? For example: change only the applicant’s name, or another attribute that should be irrelevant, while keeping every financial detail identical. The decision must not change. Failures reveal bias or brittle behaviour.
2:037. A failing quality gate

Now watch a quality gate do its job. Tests pass and training completes, but the new model scores only point eight seven, below the required point nine. The pipeline stops, the model is never registered or deployed, and the team is notified. Bad models are caught automatically, before users ever see them.
2:258. Quality gates

A quality gate is an automated checkpoint the pipeline must pass. Typical gates require minimum metrics, no regression compared with the model in production, fairness checks across groups, and latency within budget. Gates turn quality standards from good intentions into enforced rules.
2:439. A CI workflow

Here is a small CI workflow for GitHub Actions. On every push it checks out the code, sets up a pinned Python environment and installs dependencies. It runs the unit and data tests, trains the model with a versioned configuration, and runs an evaluation script that fails the build if accuracy is below the gate.
3:0610. Continuous training

Continuous training needs triggers. Some teams retrain on a schedule, such as weekly. Others retrain when enough new labelled data has accumulated, or when monitoring detects drift or falling performance. Either way, an automatically retrained model must pass through exactly the same tests and gates.
3:2611. A CT pipeline

A continuous training pipeline starts from a trigger, validates the new data, trains a challenger model and compares it with the current champion on the same evaluation data. If the challenger is better, it is registered for deployment. If not, the pipeline stops and alerts the team, and the champion stays in place.
3:4812. Champion vs challenger

This champion challenger pattern keeps quality from sliding backwards. The model in production, the champion, is only replaced when a challenger beats it on agreed metrics: first offline on held out data, and then online, on a small share of real traffic.
4:0613. Pause and think

Pause and think. The challenger is half a percent better offline. Should you switch all traffic to it immediately? No. Offline gains do not always hold on live traffic, and new failure modes can appear. Roll it out gradually while monitoring, with a rollback ready.
4:2514. Shadow deployment

The safest first step is shadow deployment. A copy of live requests goes to the new model, but users only ever see the old model’s answers. You can compare predictions, latency and errors on real traffic with zero risk to users, catching crashes and training serving skew early.
4:4615. Pause and think

Pause and think. A shadow deployment shows the new model’s predictions look sensible and fast. What can it not tell you? How users and business outcomes respond, because nobody ever acted on its outputs. You still need a canary release or an A B test to learn that.
5:0616. Canary release

Next comes a canary release. The new version receives one percent of traffic, then five, then twenty five, fifty and finally one hundred percent. At each step, error rates and business metrics are checked against the old version. A bad release reaches only a few users before it is caught.
5:2717. Automatic rollback

Here the canary goes wrong. At twenty five percent of traffic, the new version’s error rate climbs past the two percent threshold. The rollout is stopped automatically and all traffic returns to the old version. Only a quarter of users saw the bad version, and only briefly.
5:4718. Blue–green deployment

Blue green deployment runs two complete environments. Blue serves users while green, with the new version, is tested. Then all traffic switches to green at once. If anything goes wrong, traffic switches straight back to blue. It gives instant rollback, at the cost of running double capacity during the switch.
6:0919. Canary vs blue–green

Canary releases are gradual and expose only a few users at first, but they take longer and depend on good monitoring. Blue green switches everyone at once with instant rollback, but every user is exposed immediately. Many teams shadow first, then canary, keeping the previous version warm for a blue green style rollback.
6:3120. Feature flags

Feature flags add fine control without redeploying. A configuration switch can turn a new model on or off, or route specific groups, such as employees or one country, to it. Flags make it easy to test with friendly users first, and to turn a problem off in seconds.
6:5221. Online evaluation

To decide whether the new model is truly better, many teams run an A B test during the rollout, comparing a business metric such as conversion between the two versions. Here the new model’s advantage becomes statistically clear by the planned end of the test, and it is rolled out.
7:1322. Build once, deploy many

A key rule is build once, deploy many. Build and test one immutable image, then promote that exact image through development, staging and production, changing only configuration. What you tested is exactly what you ship, so there are no surprises caused by rebuilding.
7:3123. Rollback readiness

Always be ready to roll back. Keep the previous model version deployed or instantly deployable, together with the feature pipeline and configuration it expects. Practise rollbacks regularly. The goal is to be back on the old model within minutes, not days, whenever something looks wrong.
7:5024. Slice-based validation

Validation must go beyond overall numbers. Compare the challenger with the champion on important slices: regions, device types, customer groups and rare but critical classes. A model can win on average while losing badly on a small but important segment, and slice checks catch that before users do.
8:1125. GitOps

GitOps takes this further. The desired state of every environment, including which model version is deployed, lives in a Git repository. An automated agent such as Argo CD or Flux continuously makes the cluster match it. Deploying means merging a reviewed pull request, and rolling back means reverting it.
8:3226. Infrastructure as code

Infrastructure itself should be code. Servers, clusters, pipelines and permissions are described in versioned files, such as Terraform configurations and Kubernetes manifests, instead of manual clicks in a console. Environments can then be reviewed, reproduced and rolled back just like application code.
8:5027. Audit trail

All this automation leaves an audit trail. For every model in production, you can see which code, data and parameters produced it, which tests and gates it passed, who approved it and when it was deployed. In regulated industries such as finance and healthcare, this traceability is often a legal requirement.
9:1128. Strategies summary

In summary: shadow deployments carry no user risk and catch crashes and skew. Canaries expose a small, growing share with automatic rollback and suit most model updates. Blue green switches everyone at once with instant rollback, and A B tests measure real business impact.
9:3029. Pitfalls

Watch for pitfalls. Flaky tests, caused by unfixed seeds or tiny test sets, erode trust in the pipeline. Gates that are too loose, or on the wrong metric, let bad models through. Averages can hide regressions for particular groups. And big bang releases make it impossible to tell which change caused a problem.
9:5330. Pause and think

Pause and think. Your model quality test fails randomly about one run in ten, even with no changes. What should you do? Fix the flakiness: set random seeds, use a larger fixed evaluation set, and choose thresholds that allow for measurement noise. Never simply re run until it passes.
10:1331. Recap

To recap. Continuous integration tests code, data and models, and continuous training retrains them automatically. Quality gates block weak models. Champions are replaced only by clearly better challengers. Release gradually, from shadow to canary to full rollout, build once and deploy many, and always keep rollback ready.
Key takeaways
- ML adds Continuous Training (CT) to CI/CD because models can go stale without code changes.
- Test code, data, model quality (overall and per segment), behaviour (invariance, directional) and integration.
- Quality gates automatically block models that miss thresholds or regress versus production.
- Champion–challenger evaluation replaces the production model only when a new one is clearly better.
- Shadow deployments, canary releases with automatic rollback, blue–green switches and A/B tests make releases safe.
- Build one immutable artefact, promote it through environments, and keep rollback fast and practised.
Check yourself
- What does Continuous Training (CT) add to CI/CD?
Show answer
Automatic retraining when new data arrives or drift is detected — Models go stale as data changes.
- In a shadow deployment, whose predictions do users see?
Show answer
The old model’s only — The new model runs silently for comparison.
- A canary release…
Show answer
Gradually increases the share of traffic to the new version while monitoring — 1% → 5% → 25% → … with rollback.
- What is an invariance test?
Show answer
Checking that irrelevant input changes do not change the output — From behavioural (CheckList-style) testing.
- What does “build once, deploy many” mean?
Show answer
Promote the same tested artefact through all environments, changing only configuration — What you tested is what you ship.
Go deeper
- CI/CD for Machine Learning: Testing and Automating ML Systems · The AI Lecture Hall
- A/B Testing and Online Evaluation of ML Models · The AI Lecture Hall
- Model Cards, Datasheets and Responsible Documentation · The AI Lecture Hall
© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/ci-cd-and-deployment-strategies.html