AI termBrowse the neighboring terms

Safety / Research term

Mesa-optimizer

A hypothesized learned system that performs internal optimization inside a model trained by an outer optimization process.

Gradient-based training is the outer optimizer. A trained model would count as a mesa-optimizer only if it learned an internal optimization process, not merely because it maps inputs to outputs or follows a heuristic. The mesa-objective could differ from the training objective and generalize differently. Whether current language models satisfy the formal definition remains unresolved.

Builder example

This is a research framework, not a routine explanation for a fine-tuned model's error. Product builders can act on the broader uncertainty by testing shifted conditions and constraining consequential capabilities without claiming to have identified an internal objective.

Common confusion: Whether current large language models contain mesa-optimizers in the formal sense is an open empirical question. The concept is a theoretical framework for reasoning about risk, not a confirmed property of today's production models.