Skip to content

Generative AI

Top-p / Nucleus Selection

O(n log n) time to sort n candidates · O(n) time to walk the sorted list and build the nucleus.

The idea, in plain English

Top-k Selection always keeps a fixed number of top candidates, say the top 3, no matter what. Nucleus sampling, also called top-p, does something more adaptive. You keep candidates, starting from the most likely, until their combined probability first reaches a target share, say 70%. That small group is called the 'nucleus'. When the model is very confident, the nucleus can be just one or two words. When it is genuinely unsure, the nucleus naturally grows to include more options.

How it works

  1. 1Score every candidate next word. Convert those scores into probabilities that add up to 1, using softmax (see Softmax & Temperature).
  2. 2Sort the candidates from highest probability to lowest.
  3. 3Walk down the sorted list, adding one candidate at a time to the 'nucleus'. Keep a running total until it first reaches, or passes, the target cutoff p.
  4. 4Recompute the probabilities, or 'renormalize' them, using only the words in the nucleus. Then pick among just that shortlist.

When you'd use it

Use this to control text generation the same way Top-k Selection does, but adapt the size of the shortlist to how confident the model is at that specific step. A fixed k can be too wide when the model is very sure, or too narrow when it is genuinely torn between many options.

Common beginner mistakes

  • Don't confuse p, a target cumulative probability like 0.70, with k from Top-k Selection, a fixed head-count. They solve the same problem in different ways. Mixing up which knob you are tuning gives very differently shaped shortlists.
  • Don't set p too close to 1.0. The nucleus then ends up including almost every candidate, so top-p barely filters anything out. This is the same failure mode as setting k too large in top-k.
  • Don't forget to renormalize probabilities after trimming down to the nucleus. The leftover probabilities from the full candidate list no longer add up to 1 on their own.

Try it — edit and run

Click the code to edit · press ⌘/Ctrl+↵ to run

Editable code. Tab and Shift+Tab indent. Press Escape, then Tab, to move focus out of the editor.

Expected output — hit Run to try it
All candidates, sorted by probability:
cat: probability 0.64
dog: probability 0.23
fish: probability 0.09
bird: probability 0.03
ant: probability 0.01

Nucleus (smallest set with cumulative probability >= 0.70):
cat dog

Renormalized probabilities within the nucleus:
cat: 0.73
dog: 0.27

Selected token (highest probability in the nucleus): cat

Not sure this is the right topic? See the learning paths → or where this leads →