Skip to content

Data Science

One-Hot Encoding

It takes O(1) time to encode a single item once you know the category list, since that is just a length-k list of 0s and one 1. It takes O(n * k) time to encode n items across k categories.

The idea, in plain English

Computers understand numbers, not words. So if a feature is a category, like 'red', 'green', or 'blue', you cannot just hand it to a model as text. One-hot encoding turns each category into its own on/off light switch. Line up every possible category in a row. Flip on exactly one switch: the one matching this item's category. Leave every other switch off, like a row of bulbs where only one is ever lit.

How it works

  1. 1List every possible category once, in a fixed order. That fixed order becomes the position of each switch.
  2. 2For a given item, create a list of 0s the same length as the category list.
  3. 3Set a 1 at the position matching this item's category. Leave every other position at 0.

When you'd use it

Use this when preparing categorical data (colors, countries, product types) for machine learning models that expect numbers, not text, and that should not assume one category is 'bigger' or 'closer' to another.

Common beginner mistakes

  • Do not use plain numbers instead (red=1, green=2, blue=3). That accidentally tells the model blue is 'more' than red, which makes no sense for categories.
  • Do not use a different category order for different items. The same category must always land in the same switch position.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Categories: blue green red
red -> 0 0 1
blue -> 1 0 0
green -> 0 1 0
red -> 0 0 1

Not sure this is the right topic? See the learning paths → or where this leads →