Data Science
Min-Max Normalization
You take one pass to find the min and max, then one more pass to normalize — O(n) time, O(n) space for the output.
The idea, in plain English
Say one test is graded out of 50 and another out of 500. You cannot compare the raw scores fairly. Normalization fixes this by putting every score on the same 0-to-1 scale. Min-max normalization does exactly that. The smallest value in your data becomes 0. The largest becomes 1. Everything else lands proportionally in between.
How it works
- 1Find the smallest (min) and largest (max) values in your data.
- 2For each value, subtract the min, then divide by the range (max minus min).
- 3The smallest value becomes exactly 0, and the largest becomes exactly 1. Everything else falls proportionally between them.
When you'd use it
Use this before feeding numeric features into machine learning models that are sensitive to scale, like k-nearest neighbors or gradient descent. This way, a feature measured in the thousands (like salary) does not drown out one measured in single digits (like age).
Common beginner mistakes
- Do not normalize before splitting into training and test data. Doing so leaks information about the test set into training.
- Watch out for dividing by zero when every value in the data is identical (max equals min). That case needs special handling.
Try it — edit and run
Click the code to edit · press ⌘/Ctrl+↵ to run
Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.
Data: 10 20 30 40 50
Normalized: 0.00 0.25 0.50 0.75 1.00Not sure this is the right topic? See the learning paths → or where this leads →