AI in Motion

Machine LearningBeginner1:32 video6 chapters

Train/Test Splits and Cross-Validation — lecture notes

Why data is split into training, validation and test sets — and how k-fold cross-validation gives a more reliable score.

▶ Watch the animated lecture

0:001. Introduction

Introduction — Train/Test Splits and Cross-Validation

If students see the exam questions in advance, their scores mean nothing. Models are the same. To know how good a model really is, we must test it on data it has never seen.

0:142. Splitting the data

Splitting the data — Train/Test Splits and Cross-Validation

First shuffle the data so every part is representative. Then split it. Around seventy percent for training, fifteen percent for validation to tune settings, and fifteen percent for a final test that we look at only once.

0:313. Roles

Roles — Train/Test Splits and Cross-Validation

The model learns from the training set. We use the validation set to choose settings and compare models. The test set gives one final honest estimate. If you keep tuning on the test set, it stops being honest.

0:474. Cross-validation

Cross-validation — Train/Test Splits and Cross-Validation

With small datasets, one split can be lucky or unlucky. Five-fold cross-validation splits the data into five parts. Each round trains on four and validates on the fifth, until every part has been the validation set once. Here the five example scores average 0.848.

1:065. Data leakage

Data leakage — Train/Test Splits and Cross-Validation

Beware of data leakage: when information from the test data sneaks into training. Duplicates, preprocessing before splitting, features that secretly contain the answer, or shuffling time series. Leakage makes models look brilliant, then fail in reality.

1:226. Recap

Recap — Train/Test Splits and Cross-Validation

To recap. Shuffle and split. Tune on validation and test only once. Use cross-validation for small data. And guard carefully against leakage.

Key takeaways

  • Split data into training, validation and test sets (e.g. 70/15/15).
  • Tune on validation data; use the test set only once at the end.
  • k-fold cross-validation rotates the validation fold and averages the scores.
  • Data leakage makes results look better than they really are.

Check yourself

  1. Which set should you use to choose hyperparameters?
    Show answer

    Validation set — Tuning on validation keeps the test set honest.

  2. In 5-fold cross-validation, how many times is the model trained?
    Show answer

    Five times — Each of the five folds takes a turn as validation data.

  3. Which is an example of data leakage?
    Show answer

    Normalising the full dataset before splitting it — Test statistics influence the training data when preprocessing happens before the split.

Go deeper

© 2026 Janin A Apurba, CSE, AUST · Advanced ICT Officer, CNRS-UNHCR. All rights reserved. Notes for the animated lecture at https://ai-in-motion.vercel.app/watch/train-test-split-and-cross-validation.html