AI in Motion

MLOps & EngineeringDeep diveIntermediate10:33 video31 chapters

CI/CD for ML and Safe Deployment Strategies — lecture notes

Automate the path to production: testing code, data and models; CI/CD/CT pipelines with quality gates; champion–challenger evaluation; shadow, canary and blue–green deployments; and fast rollback.

▶ Watch the animated lecture

0:001. Introduction

Introduction — CI/CD for ML and Safe Deployment Strategies

How do you change a model that millions of people rely on without breaking anything? You automate the path to production and release changes gradually, with the ability to undo them instantly. This deep dive covers continuous integration, delivery and training for machine learning, and the deployment strategies that keep releases safe.

0:212. CI, CD and CT

CI, CD and CT — CI/CD for ML and Safe Deployment Strategies

Three ideas work together. Continuous integration tests every change automatically. Continuous delivery, or deployment, releases changes automatically and safely. Machine learning adds continuous training, which retrains models automatically when new data arrives, because a model can go stale even when the code does not change.

0:413. A CI/CD pipeline

A CI/CD pipeline — CI/CD for ML and Safe Deployment Strategies

Here is a pipeline in action. A commit triggers unit tests and a data schema check. The model is trained on the latest data, then evaluated: accuracy must reach at least point nine and there must be no fairness regression. The model is registered, and finally deployed as a canary to five percent of traffic.

1:044. Tests for ML

Tests for ML — CI/CD for ML and Safe Deployment Strategies

Machine learning needs several kinds of tests. Unit tests check code logic. Data tests check schemas, ranges and freshness. Model quality tests compare metrics with thresholds, overall and for important segments. Behavioural tests check expected behaviours, and integration tests check the whole service end to end.

1:245. Behavioural testing

Behavioural testing — CI/CD for ML and Safe Deployment Strategies

Behavioural testing, popularised by the CheckList method, treats a model like software with specifications. Invariance tests check that irrelevant changes, such as swapping a person’s name, do not change the output. Directional tests check that some changes move the output the expected way, like adding the word not.

1:446. Pause and think

Pause and think — CI/CD for ML and Safe Deployment Strategies

Pause and think. Can you write an invariance test for a loan approval model? For example: change only the applicant’s name, or another attribute that should be irrelevant, while keeping every financial detail identical. The decision must not change. Failures reveal bias or brittle behaviour.

2:037. A failing quality gate

A failing quality gate — CI/CD for ML and Safe Deployment Strategies

Now watch a quality gate do its job. Tests pass and training completes, but the new model scores only point eight seven, below the required point nine. The pipeline stops, the model is never registered or deployed, and the team is notified. Bad models are caught automatically, before users ever see them.

2:258. Quality gates

Quality gates — CI/CD for ML and Safe Deployment Strategies

A quality gate is an automated checkpoint the pipeline must pass. Typical gates require minimum metrics, no regression compared with the model in production, fairness checks across groups, and latency within budget. Gates turn quality standards from good intentions into enforced rules.

2:439. A CI workflow

A CI workflow — CI/CD for ML and Safe Deployment Strategies

Here is a small CI workflow for GitHub Actions. On every push it checks out the code, sets up a pinned Python environment and installs dependencies. It runs the unit and data tests, trains the model with a versioned configuration, and runs an evaluation script that fails the build if accuracy is below the gate.

3:0610. Continuous training

Continuous training — CI/CD for ML and Safe Deployment Strategies

Continuous training needs triggers. Some teams retrain on a schedule, such as weekly. Others retrain when enough new labelled data has accumulated, or when monitoring detects drift or falling performance. Either way, an automatically retrained model must pass through exactly the same tests and gates.

3:2611. A CT pipeline

A CT pipeline — CI/CD for ML and Safe Deployment Strategies

A continuous training pipeline starts from a trigger, validates the new data, trains a challenger model and compares it with the current champion on the same evaluation data. If the challenger is better, it is registered for deployment. If not, the pipeline stops and alerts the team, and the champion stays in place.

3:4812. Champion vs challenger

Champion vs challenger — CI/CD for ML and Safe Deployment Strategies

This champion challenger pattern keeps quality from sliding backwards. The model in production, the champion, is only replaced when a challenger beats it on agreed metrics: first offline on held out data, and then online, on a small share of real traffic.

4:0613. Pause and think

Pause and think — CI/CD for ML and Safe Deployment Strategies

Pause and think. The challenger is half a percent better offline. Should you switch all traffic to it immediately? No. Offline gains do not always hold on live traffic, and new failure modes can appear. Roll it out gradually while monitoring, with a rollback ready.

4:2514. Shadow deployment

Shadow deployment — CI/CD for ML and Safe Deployment Strategies

The safest first step is shadow deployment. A copy of live requests goes to the new model, but users only ever see the old model’s answers. You can compare predictions, latency and errors on real traffic with zero risk to users, catching crashes and training serving skew early.

4:4615. Pause and think

Pause and think — CI/CD for ML and Safe Deployment Strategies

Pause and think. A shadow deployment shows the new model’s predictions look sensible and fast. What can it not tell you? How users and business outcomes respond, because nobody ever acted on its outputs. You still need a canary release or an A B test to learn that.

5:0616. Canary release

Canary release — CI/CD for ML and Safe Deployment Strategies

Next comes a canary release. The new version receives one percent of traffic, then five, then twenty five, fifty and finally one hundred percent. At each step, error rates and business metrics are checked against the old version. A bad release reaches only a few users before it is caught.

5:2717. Automatic rollback

Automatic rollback — CI/CD for ML and Safe Deployment Strategies

Here the canary goes wrong. At twenty five percent of traffic, the new version’s error rate climbs past the two percent threshold. The rollout is stopped automatically and all traffic returns to the old version. Only a quarter of users saw the bad version, and only briefly.

5:4718. Blue–green deployment

Blue–green deployment — CI/CD for ML and Safe Deployment Strategies

Blue green deployment runs two complete environments. Blue serves users while green, with the new version, is tested. Then all traffic switches to green at once. If anything goes wrong, traffic switches straight back to blue. It gives instant rollback, at the cost of running double capacity during the switch.

6:0919. Canary vs blue–green

Canary vs blue–green — CI/CD for ML and Safe Deployment Strategies

Canary releases are gradual and expose only a few users at first, but they take longer and depend on good monitoring. Blue green switches everyone at once with instant rollback, but every user is exposed immediately. Many teams shadow first, then canary, keeping the previous version warm for a blue green style rollback.

6:3120. Feature flags

Feature flags — CI/CD for ML and Safe Deployment Strategies

Feature flags add fine control without redeploying. A configuration switch can turn a new model on or off, or route specific groups, such as employees or one country, to it. Flags make it easy to test with friendly users first, and to turn a problem off in seconds.

6:5221. Online evaluation

Online evaluation — CI/CD for ML and Safe Deployment Strategies

To decide whether the new model is truly better, many teams run an A B test during the rollout, comparing a business metric such as conversion between the two versions. Here the new model’s advantage becomes statistically clear by the planned end of the test, and it is rolled out.

7:1322. Build once, deploy many

Build once, deploy many — CI/CD for ML and Safe Deployment Strategies

A key rule is build once, deploy many. Build and test one immutable image, then promote that exact image through development, staging and production, changing only configuration. What you tested is exactly what you ship, so there are no surprises caused by rebuilding.

7:3123. Rollback readiness

Rollback readiness — CI/CD for ML and Safe Deployment Strategies

Always be ready to roll back. Keep the previous model version deployed or instantly deployable, together with the feature pipeline and configuration it expects. Practise rollbacks regularly. The goal is to be back on the old model within minutes, not days, whenever something looks wrong.

7:5024. Slice-based validation

Slice-based validation — CI/CD for ML and Safe Deployment Strategies

Validation must go beyond overall numbers. Compare the challenger with the champion on important slices: regions, device types, customer groups and rare but critical classes. A model can win on average while losing badly on a small but important segment, and slice checks catch that before users do.

8:1125. GitOps

GitOps — CI/CD for ML and Safe Deployment Strategies

GitOps takes this further. The desired state of every environment, including which model version is deployed, lives in a Git repository. An automated agent such as Argo CD or Flux continuously makes the cluster match it. Deploying means merging a reviewed pull request, and rolling back means reverting it.

8:3226. Infrastructure as code

Infrastructure as code — CI/CD for ML and Safe Deployment Strategies

Infrastructure itself should be code. Servers, clusters, pipelines and permissions are described in versioned files, such as Terraform configurations and Kubernetes manifests, instead of manual clicks in a console. Environments can then be reviewed, reproduced and rolled back just like application code.

8:5027. Audit trail

Audit trail — CI/CD for ML and Safe Deployment Strategies

All this automation leaves an audit trail. For every model in production, you can see which code, data and parameters produced it, which tests and gates it passed, who approved it and when it was deployed. In regulated industries such as finance and healthcare, this traceability is often a legal requirement.

9:1128. Strategies summary

Strategies summary — CI/CD for ML and Safe Deployment Strategies

In summary: shadow deployments carry no user risk and catch crashes and skew. Canaries expose a small, growing share with automatic rollback and suit most model updates. Blue green switches everyone at once with instant rollback, and A B tests measure real business impact.

9:3029. Pitfalls

Pitfalls — CI/CD for ML and Safe Deployment Strategies

Watch for pitfalls. Flaky tests, caused by unfixed seeds or tiny test sets, erode trust in the pipeline. Gates that are too loose, or on the wrong metric, let bad models through. Averages can hide regressions for particular groups. And big bang releases make it impossible to tell which change caused a problem.

9:5330. Pause and think

Pause and think — CI/CD for ML and Safe Deployment Strategies

Pause and think. Your model quality test fails randomly about one run in ten, even with no changes. What should you do? Fix the flakiness: set random seeds, use a larger fixed evaluation set, and choose thresholds that allow for measurement noise. Never simply re run until it passes.

10:1331. Recap

Recap — CI/CD for ML and Safe Deployment Strategies

To recap. Continuous integration tests code, data and models, and continuous training retrains them automatically. Quality gates block weak models. Champions are replaced only by clearly better challengers. Release gradually, from shadow to canary to full rollout, build once and deploy many, and always keep rollback ready.

Key takeaways

  • ML adds Continuous Training (CT) to CI/CD because models can go stale without code changes.
  • Test code, data, model quality (overall and per segment), behaviour (invariance, directional) and integration.
  • Quality gates automatically block models that miss thresholds or regress versus production.
  • Champion–challenger evaluation replaces the production model only when a new one is clearly better.
  • Shadow deployments, canary releases with automatic rollback, blue–green switches and A/B tests make releases safe.
  • Build one immutable artefact, promote it through environments, and keep rollback fast and practised.

Check yourself

  1. What does Continuous Training (CT) add to CI/CD?
    Show answer

    Automatic retraining when new data arrives or drift is detected — Models go stale as data changes.

  2. In a shadow deployment, whose predictions do users see?
    Show answer

    The old model’s only — The new model runs silently for comparison.

  3. A canary release…
    Show answer

    Gradually increases the share of traffic to the new version while monitoring — 1% → 5% → 25% → … with rollback.

  4. What is an invariance test?
    Show answer

    Checking that irrelevant input changes do not change the output — From behavioural (CheckList-style) testing.

  5. What does “build once, deploy many” mean?
    Show answer

    Promote the same tested artefact through all environments, changing only configuration — What you tested is what you ship.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/ci-cd-and-deployment-strategies.html