Skip to content

Generative AI

Perplexity

O(n) time, where n is the number of words being scored.

The idea, in plain English

Imagine you read a sentence out loud with a friend. Every time you say the next word, they silently guess whether they saw it coming. A friend who is rarely surprised, who keeps nodding 'yeah, I expected that', knows you well. Perplexity turns that idea into one number for a language model. It measures how surprised the model was, on average, by the words that actually came next in some real text. A model that gives high probability to what actually happens is rarely surprised, and gets a LOW perplexity score. A model that keeps getting caught off guard gets a HIGH perplexity score.

How it works

  1. 1For each word in a real piece of text, take the probability the model gave to that exact word being next. In a real system, this comes from softmax over the whole vocabulary at that step.
  2. 2Multiply all of these per-word probabilities together. That gives you the probability the model assigned to the entire sequence happening exactly as it did.
  3. 3This combined probability shrinks fast as sentences get longer. To make sequences of different lengths comparable, invert it (1 divided by it) and take the n-th root, where n is the number of words. The result is perplexity: a per-word score for how surprised the model was, on average.

When you'd use it

This is the standard headline metric for comparing how well two language models predict real text, or for tracking whether a model is improving during training. Lower is always better, whatever the model's exact architecture.

Common beginner mistakes

  • Don't read perplexity backwards. It is easy to instinctively assume a higher score is better, but perplexity measures confusion, so a lower score always means a better fit to the real text.
  • Don't compare perplexity scores that were computed differently, such as over different vocabularies, different tokenization, or different text. Perplexity is only a fair comparison between models scored the exact same way on the exact same text.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
Real sentence: the cat sat

Model A (confident):
  probability assigned to "the": 0.50
  probability assigned to "cat": 0.50
  probability assigned to "sat": 0.50
  perplexity: 2.00

Model B (unsure):
  probability assigned to "the": 0.20
  probability assigned to "cat": 0.20
  probability assigned to "sat": 0.20
  perplexity: 5.00

Lower perplexity means less surprised: Model A (confident) fits the real sentence better.

Not sure this is the right topic? See the learning paths → or where this leads →