AI termBrowse the neighboring terms

Training / Standard term

Reinforcement Learning with Verifiable Rewards (RLVR)

Reinforcement learning that uses programmatically checkable reward signals such as exact answers, executable tests, or formal verification.

RLVR replaces some learned or subjective reward judgments with checks whose output can be computed. The checker is still a specification written by people and may be incomplete, exploitable, or only loosely connected to the real goal. A passing unit test establishes the tested behavior; it does not make all code quality or security objectively settled.

Builder example

Verifiable tasks can supply large amounts of low-cost feedback and support search or training. Results vary with the base model, reward coverage, optimization method, and task. Open-ended work can also contain verifiable subproblems without reducing the whole objective to one score.

Common confusion: Verifiable does not mean value-free or complete. The check can be deterministic while the choice of what to check remains a design judgment.