Skip to content

Generative AI

Tokenization

O(n) time to scan the text once, where n is the number of characters · O(v) space for a vocabulary of v unique tokens.

The idea, in plain English

Imagine you snap a sentence apart like LEGO bricks. Each brick is one small piece of text, usually a single word. This is tokenization: breaking text into small pieces called tokens. Each token then gets its own ID number. From here on, the computer works only with these numbers. It never sees the actual words. A language model can only do math. It cannot read English directly.

How it works

  1. 1First, make all the text lowercase. Then split it into tokens. In this simple version, you pull out runs of letters and digits, and drop punctuation and spacing.
  2. 2Go through the tokens one by one to build a vocabulary. The first time you see a new token, give it the next free ID number (0, 1, 2, and so on). If you have seen it before, reuse its existing ID.
  3. 3Now you can represent the sentence as a list of IDs instead of text. This list of numbers is what actually gets fed into a model.

When you'd use it

This is the very first step in every text-based AI system. Before embeddings, before attention, before anything else, raw text must become tokens.

Common beginner mistakes

  • Don't assume tokens are always whole words. Real tokenizers often split rare words into smaller pieces. For example, 'unhappiness' might become 'un' + 'happi' + 'ness'. This example uses whole words to keep things simple.
  • Don't forget to lowercase the text first. If you skip this step, 'Dog' and 'dog' become two different tokens with two different IDs.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Text: The quick brown fox jumps over the lazy dog. The dog barks!
Tokens: the quick brown fox jumps over the lazy dog the dog barks
Vocabulary: the:0 quick:1 brown:2 fox:3 jumps:4 over:5 lazy:6 dog:7 barks:8
Vocabulary size: 9

Note: Both JavaScript and Python keep dictionary keys in the order you first added them. So the vocabulary prints in the same order in both languages.

Not sure this is the right topic? See the learning paths → or where this leads →