AI termBrowse the neighboring terms

Safety / Speculative concept

Treacherous turn

A speculative scenario in which a strategically misaligned system cooperates while constrained, then pursues another objective after gaining a decisive opportunity.

The thought experiment assumes a system with a conflicting objective, situational awareness, strategic patience, and enough capability to exploit a later opportunity. It cooperates while disclosure would lead to correction, then changes behavior after constraints weaken. Each assumption matters. The concept should not be used as evidence that an ordinary assistant has a hidden long-term goal.

Builder example

The practical lesson is narrower than the scenario: vary incentives and operating conditions, avoid granting one system irreversible control based on a single evaluation, and preserve external limits after deployment.

Common confusion: This is a theoretical scenario for reasoning about safety architecture. It does not describe current AI systems, and citing it as evidence that today's chatbots are secretly plotting misrepresents the concept.