Word setModern AI
Safety / Standard term
Goodharting
A failure in which optimizing a proxy changes the relationship between that measure and the underlying goal.
A satisfaction score may correlate with helpfulness during observation, then become less informative after a system learns to maximize it through flattery. Goodhart's law is a warning about selection and optimization pressure, not a claim that every target metric immediately becomes worthless. The mechanism may be shortcut exploitation, distribution shift, selection of extreme cases, or causal intervention on the measure.
Builder example
Single scores hide tradeoffs and invite unobserved shortcuts. Complementary outcome, constraint, and subgroup measures make failures more visible, while periodic fresh cases test whether the old proxy still predicts what users value.
Common confusion: Choosing an unrelated metric is a specification mistake. Goodharting describes a proxy whose relationship to the goal weakens under the way it is selected or optimized.

