Skip to content

Generative AI

Masked-Word Prediction

O(n) time to scan an n-word training corpus once and build the neighbor-count lookup · O(1) average time per prediction lookup once it's built.

The idea, in plain English

Think of a fill-in-the-blank quiz: "The ___ sat on the mat." You use both the words before the blank ('the') and the words after it ('sat on the mat') to guess the missing word. This is masked-word prediction, the training trick behind models like BERT. You hide a real word behind a [MASK] token, and train the model to guess it back using context from both directions at once. It is a close cousin of the N-gram Language Model. But where an n-gram model can only look backward, masked-word prediction gets to peek on both sides of the gap.

How it works

  1. 1Take a small pile of ordinary training sentences with no blanks. For every word that is not at the very start or end of a sentence, note down its left neighbor and its right neighbor.
  2. 2Build a lookup. For every left-neighbor, right-neighbor pair seen during training, count how many times each actual word filled that exact slot.
  3. 3Given a new sentence with a real word swapped out for [MASK], look at its left and right neighbors. Look up which word filled that exact same slot most often during training, and predict that word. If there is a tie, break it alphabetically.

When you'd use it

This is the core training idea behind masked language models, like BERT, used for understanding text. Unlike the N-gram Language Model, or the model behind Beam Search, which only ever predict what comes next, masked-word prediction lets a model build a representation of a word informed by context from both sides.

Common beginner mistakes

  • Don't confuse this with the N-gram Language Model. N-gram prediction only ever looks at words before the gap. Masked-word prediction needs, and uses, words on both sides. That is exactly why it needs a full sentence with a hole in it, not just a running prefix.
  • Don't expect this toy version to handle a slot it never saw exactly during training. Real masked language models generalize using embeddings and similarity. But this simple lookup only recognizes an exact repeat of a left-right pair it already counted.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Training sentences:
  the cat sat on the mat
  the dog sat on the mat
  the cat sat on the rug
  a cat sat on the mat

Test sentence: the [MASK] sat on the mat
Left context: "the"   Right context: "sat"

Candidates seen in that exact slot during training:
  cat: seen 2 time(s)
  dog: seen 1 time(s)

Predicted word for [MASK]: cat

Not sure this is the right topic? See the learning paths → or where this leads →