AI termBrowse the neighboring terms

Training / Standard term

Constitutional AI

A post-training approach that uses a written set of principles to generate critiques, revisions, and AI feedback for model behavior.

In Anthropic's published method, a model first critiques and revises responses under constitutional principles, producing supervised training data. A later reinforcement-learning stage uses AI-generated preference judgments guided by those principles. People still choose the constitution, training setup, and evaluations, and implementations described with the label may differ from the original pipeline.

Builder example

A written constitution makes intended principles easier to inspect than an unnamed preference target. It does not make each model decision traceable to one clause or guarantee that training generalizes as intended. Product teams can use explicit policies for evaluation even without training a model.

Common confusion: A constitution does not guarantee ethical behavior. The principles are only as good as the people who wrote them, and the training process can still produce unexpected gaps.