Skip to content

Data Science

Train/Test Split

Slicing takes O(1) time, since cutting a list into two pieces just copies each element once, with no extra scanning needed.

The idea, in plain English

Before trusting a model, you need to check its homework on questions it has never seen. That is why you split data into two piles: a training set the model learns from, and a test set you hold back to grade it honestly afterward. Here we use a simple, fixed split. We always take the same first chunk as training and the rest as testing, so the result is exactly the same every time you run it. In real projects, you would usually shuffle first with a fixed random seed, so the training set is not just 'whatever came first' in the file.

How it works

  1. 1Decide how many examples should go into the training set — for example, 7 out of 10.
  2. 2Take that many examples from the front of the data, in the order they appear, as the training set.
  3. 3Everything left over becomes the test set.

When you'd use it

Do this every time you build a predictive model. Without a held-out test set, you have no honest way to know if the model actually learned the pattern or just memorized the training data.

Common beginner mistakes

  • Do not test on the same data you trained on. That always looks great and tells you nothing about how the model handles new data.
  • Do not use a plain first-chunk split on data that is sorted or grouped by category. You could end up with a training set that never sees an entire category. Shuffle first, with a fixed seed, unless the data is already in random order.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Data: 1 2 3 4 5 6 7 8 9 10
Train size: 7
Test size: 3
Train: 1 2 3 4 5 6 7
Test: 8 9 10

Not sure this is the right topic? See the learning paths → or where this leads →