Data Pipelines, Versioning and Experiment Tracking
Treat data like code: validation, versioning with DVC, leakage-safe splits, feature stores and point-in-time correctness, experiment tracking with MLflow, reproducibility and model registries.
📄 Illustrated notes · every chapter as a picture · printable
Quick quiz
5 questions to check your understanding.
Go deeper
University-level written lectures in The AI Lecture Hall:
Transcript
Introduction. In machine learning, data is as important as code, often more so. Yet data is frequently copied around by hand, and experiments are remembered rather than recorded. This deep dive covers how to build reliable data pipelines, version datasets, avoid leakage, and track experiments so results can be trusted and reproduced.
Data-centric AI. Data centric AI is the idea that systematically improving the data, its labels, coverage and consistency, often beats endlessly tweaking the model. Fixing a few percent of mislabelled examples, or collecting examples of rare cases, can improve results more than switching to a bigger architecture.
A data pipeline. A production data pipeline has clear stages. Ingest data from databases, event streams and files. Validate it before anything else. Clean it and compute features. Split it into training, validation and test sets without leakage. Finally store immutable, versioned snapshots, so every model can point to exactly the data it used.
Data validation. Data validation is automated testing for data. Before training, and before serving, check the schema, types, allowed values, ranges, missing value rates, uniqueness and freshness. Tools like Great Expectations and TensorFlow Data Validation make these checks easy to write and run in every pipeline.
Pause and think. Pause and think. An upstream system silently switches a distance column from kilometres to miles. How could your pipeline notice? Values suddenly shrink by about thirty eight percent, so range checks and distribution comparisons with the previous batch would fire, before the bad data reaches training or serving.
Validation in code. Expectations can be written as code. With the pandera library, declare each column’s type, allowed range, permitted values and whether missing values are allowed. Validating a batch raises an error that lists every violation, so the pipeline can stop before bad data reaches training or serving.
Data contracts. Many data problems are really communication problems. A data contract is an agreed, versioned specification between the team that produces a dataset and the teams that consume it: its schema, the meaning of each field, freshness guarantees and how changes are announced. No more silent renames.
Data versioning. Code lives in Git, but large datasets do not fit there. Data versioning tools such as DVC store the actual files in cloud storage, and commit small pointer files to Git. Every commit then pins an exact version of both the code and the data, and anyone can check out and reproduce it.
Versioning with DVC. Here is the workflow. Initialise Git and DVC and configure remote storage. Adding a dataset makes DVC hash it and create a small pointer file, which you commit to Git. Push uploads the data itself. Later, checking out an old commit and pulling restores exactly the data that version used.
Lineage. With code and data both versioned, lineage connects them. Each experiment run links to a code commit and a data version, and each registered model links to its run. When a model misbehaves, you can trace back to exactly what produced it, and reproduce or fix it with confidence.
Data leakage. The most dangerous data bug is leakage: information that would not be available at prediction time sneaks into training. Examples include normalising with statistics from the full dataset, including the test set, or using a feature recorded after the outcome. Leakage makes evaluation look wonderful and production disappointing.
Leakage-safe splits. Correct splitting is the first defence. Hold out a test set that is never touched during development, and use validation data, or k fold cross validation, for choosing models. Fit every preprocessing step on training data only, then apply it to validation and test data.
Pause and think. Pause and think. You are predicting next week’s sales, and you split your time series randomly into training and test sets. What is wrong? The model trains on days after some of its test days, effectively peeking into the future. Use time based splits: train on the past, test on later periods.
Stratified and group splits. Two more splitting rules matter. Stratify splits so that rare classes appear in the same proportion everywhere. And split by group, so that all records belonging to one user, patient or device land in the same split. Otherwise the model can memorise individuals and look better than it really is.
Pause and think. Pause and think. A medical dataset contains several scans per patient, and you split the scans randomly. What goes wrong? Scans of the same patient end up in both training and test sets, so the model can recognise the patient instead of the disease, inflating test accuracy. Split by patient.
Labels. Labels deserve the same care. Write clear labelling guidelines, have several annotators label a sample, measure their agreement, and review disagreements. Low agreement usually signals an unclear task definition rather than careless people, and fixing the definition improves both the data and the model.
Feature stores. A feature store attacks training serving skew directly. Features are defined once and computed by one pipeline. Their full history goes to an offline store for training, and their latest values are synced to a fast online store. The model server reads the same features, in milliseconds, that the model was trained on.
Point-in-time correctness. Feature stores also enforce point in time correctness. When building a training example for a past date, each feature must use only values known at that moment. Counting a customer’s complaints including those filed after the prediction date is a subtle, and very common, form of leakage.
Experiment tracking. Experiment tracking replaces memory and spreadsheets. Every training run automatically records its parameters, metrics, code and data versions, environment and output files in a searchable store. Months later you can answer: which settings produced our best model, and on which data?
Comparing runs. Here five runs are tracked. Run one, with a high learning rate, peaks early and then overfits. Run three, with a tiny learning rate, learns slowly. Run four, with dropout, reaches the best validation accuracy. With every run recorded, choosing and reproducing the winner is straightforward.
Tracking with MLflow. With MLflow this takes a few lines. Name an experiment and start a run. Log the hyperparameters, train, and log the validation metric. Tag the run with the data version, and save the model itself as an artefact, ready to be registered and deployed.
What to log. Every run should record its parameters and configuration, its metrics, the Git commit of the code, the data version, the software environment and hardware, and its artefacts such as the model file and diagnostic plots. Together these make any result explainable and repeatable.
Hyperparameter search. Tracking pays off during hyperparameter search. Grid search tries every combination and wastes effort on unimportant settings. Random search samples combinations and often finds good settings faster, because usually only a few hyperparameters really matter. Bayesian optimisation learns which regions look promising.
Reproducibility. Reproducibility needs several habits: fix random seeds, pin library versions, run inside a container, version the data, log the configuration, and be aware that some GPU operations are not deterministic. The goal is that re running a pipeline gives the same metrics, within a small amount of noise.
Pause and think. Pause and think. You re run last month’s training script and get two percent lower accuracy. What could explain it? The underlying data may have changed, library versions may differ, or randomness from seeds, data order or non deterministic GPU kernels was not controlled. Versioning and pinning would tell you which.
Model registry. A model registry is the catalogue of trained models. Each version carries its metadata, lineage and a lifecycle stage, such as staging, production or archived. Deployments reference a named, versioned model rather than a loose file on someone’s laptop, which makes promotions and rollbacks clean.
Promotion flow. The flow is simple. The best tracked run is registered as a new version in staging. Automated tests, fairness checks and a human review decide whether it is promoted to production. The previous version is archived, but kept ready, so a rollback takes seconds.
Datasheets. Datasets deserve documentation too. Datasheets for datasets, proposed in 2018, record why a dataset was created, what it contains, how it was collected and labelled, its known gaps and biases, and how it may be used. They prevent a dataset from being misused long after its creators have moved on.
Privacy and governance. Data governance is not optional. Collect only what the model needs. Mask or remove personal information. Make sure you have consent and the licences to use the data. And delete data once it is no longer needed. Regulations such as the GDPR make these obligations legal requirements in many countries.
Pause and think. Pause and think. Two teams compute customer tenure differently, one in days and one in rounded months. What is the lasting fix? One shared, versioned feature definition, in a feature store or shared library, used by training, by serving and by every team.
Best practices. In short: validate data at every boundary, version datasets and link them to runs, split by time or group to avoid leakage, track every experiment automatically, and register models instead of deploying loose files. These habits make results trustworthy and teams faster.
Recap. To recap. Better data often beats a better model. Validate, version and document your data. Prevent leakage with careful splits and point in time correct features. Track every experiment, and promote models through a registry with clear stages.