Safety / Research term
Mesa-optimizer
A hypothesized learned system that performs internal optimization inside a model trained by an outer optimization process.
Gradient-based training is the outer optimizer. A trained model would count as a mesa-optimizer only if it learned an internal optimization process, not merely because it maps inputs to outputs or follows a heuristic. The mesa-objective could differ from the training objective and generalize differently. Whether current language models satisfy the formal definition remains unresolved.
Builder example
This is a research framework, not a routine explanation for a fine-tuned model's error. Product builders can act on the broader uncertainty by testing shifted conditions and constraining consequential capabilities without claiming to have identified an internal objective.
Common confusion: Whether current large language models contain mesa-optimizers in the formal sense is an open empirical question. The concept is a theoretical framework for reasoning about risk, not a confirmed property of today's production models.

