Training / Standard term
Reinforcement Learning from AI Feedback (RLAIF)
Using judgments generated by another AI system as part of the reward or preference signal for reinforcement learning.
An AI evaluator can compare responses under a rubric or written principles, producing preference data that supplements or replaces some direct human labels. People still choose the constitution, prompts, sampling, model, and validation process. Cost and scale advantages depend on the judge and pipeline rather than following automatically from the acronym.
Builder example
Judge errors and preferences can propagate into the trained policy. Documentation should describe the principles, evaluator, human oversight of the data process, and behavioral evaluations. The method name alone does not explain a deployed model's tone or refusal pattern.
Common confusion: AI feedback is not inherently more objective than human feedback. The judge model carries its own biases and blind spots, which get baked into the training signal just as human biases would.

