Data Science
One-Hot Encoding
It takes O(1) time to encode a single item once you know the category list, since that is just a length-k list of 0s and one 1. It takes O(n * k) time to encode n items across k categories.
The idea, in plain English
Computers understand numbers, not words. So if a feature is a category, like 'red', 'green', or 'blue', you cannot just hand it to a model as text. One-hot encoding turns each category into its own on/off light switch. Line up every possible category in a row. Flip on exactly one switch: the one matching this item's category. Leave every other switch off, like a row of bulbs where only one is ever lit.
How it works
- 1List every possible category once, in a fixed order. That fixed order becomes the position of each switch.
- 2For a given item, create a list of 0s the same length as the category list.
- 3Set a 1 at the position matching this item's category. Leave every other position at 0.
When you'd use it
Use this when preparing categorical data (colors, countries, product types) for machine learning models that expect numbers, not text, and that should not assume one category is 'bigger' or 'closer' to another.
Common beginner mistakes
- Do not use plain numbers instead (red=1, green=2, blue=3). That accidentally tells the model blue is 'more' than red, which makes no sense for categories.
- Do not use a different category order for different items. The same category must always land in the same switch position.
Try it — edit and run
Click the code to edit · press ⌘/Ctrl+↵ to run
Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.
Categories: blue green red
red -> 0 0 1
blue -> 1 0 0
green -> 0 1 0
red -> 0 0 1Not sure this is the right topic? See the learning paths → or where this leads →