Generative AI
N-gram Language Model
O(n) time to build the counts from n training tokens · O(1) average time per prediction lookup once built.
The idea, in plain English
When your phone keyboard suggests the next word as you type, it is doing something close to this old, simple trick: just count. Before today's huge AI models existed, you could predict the next word by reading a big pile of text. For every word, you tally which word tends to follow it directly. To predict what comes next after a word, you pick whichever follower showed up most often in your counts. This is the ancestor of your keyboard's suggestions, and of what a modern language model does, just at a much bigger scale.
How it works
- 1Break the training text into a sequence of tokens, or words.
- 2For every word, count how often each other word directly follows it. These word pairs are called 'bigrams', or 2-grams.
- 3To predict the next word after a word W, look up W's followers and pick the one with the highest count. If there is a tie, break it alphabetically, so the result stays predictable.
When you'd use it
This is a lightweight, fully explainable way to model what word comes next. It is useful for teaching the core idea before you move to neural language models, or for simple autocomplete.
Common beginner mistakes
- Don't rely on looking only one word back, which is called a 'bigram'. Real language depends on much more context. That is why modern models look back over huge windows of text.
- Watch out for words the model never saw during training. An n-gram model has no prediction for a word it has no counts for.
Try it — edit and run
Click the code to edit · press ⌘/Ctrl+↵ to run
Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.
Corpus: the cat sat on the mat the cat ran on the rug the dog sat on the mat
Prediction after 'the': cat
Prediction after 'cat': ran
Prediction after 'on': the
Prediction after 'sat': onNote: 'The' has two equally common followers, 'cat' and 'mat', each seen twice. 'Cat' also has two equally common followers, 'ran' and 'sat', each seen once. Ties are broken alphabetically, so both languages always pick the same word.
Not sure this is the right topic? See the learning paths → or where this leads →