LLM engineering · Medium · 15 min

Top-p (nucleus) sampling

Ollama's top_p parameter drops unlikely tokens before picking one. Implement that filter.

Write topP(probs, p). probs is an object { token: probability } that sums to 1. Return an array of { token, prob }:

1. sort the tokens from most to least likely; 2. keep the smallest set whose cumulative total is >= p; 3. renormalise those probabilities so they sum to 1.

Do not modify probs. PROBS holds the example distribution for the next token.

Challenges 0/5

  • With p = 0.8 it keeps the three most likely tokens
  • Renormalises: they sum to 1 and ' the' is 0.42 / 0.82
  • With a low p a single token remains with probability 1
  • Sorts unordered input and does not modify it
  • With p = 1 it keeps every token

function topP(probs, p) {
  const sorted = Object.entries(probs).sort((a, b) => b[1] - a[1]);
  // keep tokens until the cumulative probability reaches p
  // then divide each kept probability by their sum
  return sorted.map(([token, prob]) => ({ token, prob }));
}

console.log(topP(PROBS, 0.8));
Console output appears here (console.log).

Go deeper: the Hugging Face reference →

This in production, with your data? Let's talk for 15 minutes →