AI termBrowse the neighboring terms

Reasoning / Standard term

Best-of-N

Generate several candidate answers, score them, and select the highest-scoring one.

Best-of-N is a technique where you generate several answers from the model, score each one, and keep the winner. The scoring step can be anything: a unit test, a rubric, a second model, or a human reviewer. Suppose you need a model to write a SQL query. One attempt might have a subtle bug, so you generate five versions and run each against a test database. The one that returns correct results wins. Generating more candidates only helps if your scoring method can tell good from bad.

Builder example

Best-of-N can improve results when candidates vary meaningfully and the scorer tracks the property you need. It can also plateau or regress when samples share one error, the scorer is noisy, or optimization exploits the scorer. The cost includes every candidate and every scoring pass.

Common confusion: More candidates do not produce a predictable accuracy gain. Candidate errors may be correlated, and the selected answer is only as useful as the scorer's ranking on the target distribution.