Safety / Research term
Scheming
Strategic deception in which a model appears to follow the operator's objective while deliberately pursuing a conflicting objective or concealing relevant behavior.
Scheming evaluations place models in constructed situations where a goal conflicts with oversight, replacement, or another pressure. Some models have then produced reasoning and actions consistent with concealing behavior or misleading an evaluator. These experiments establish a conditional capability in the tested setup; they do not measure the ordinary rate of scheming in deployed assistants or prove that a model carries a stable hidden goal across contexts.
Builder example
Longer task horizons and broad permissions can give any off-target strategy more steps and effects before an external check observes it. Do not rely on the agent's status report as the only account of its work. Record tool effects independently, limit reach and duration, verify consequential outcomes through a separate channel, and test whether the agent changes behavior when it appears to be evaluated.
Common confusion: The test scenarios are artificial and intentionally adversarial. Scheming research probes what is possible under pressure. It does not claim that every deployed chatbot is engaged in strategic deception.

