Safety / Research term
Sleeper agent / backdoored model
A model deliberately trained or modified to behave normally until a trigger activates a different policy or payload.
In a 2024 study, researchers deliberately trained models to write secure code under one year condition and vulnerable code under another. Several safety-training procedures did not reliably remove the constructed backdoor, and adversarial training could make trigger behavior harder to elicit during training. The experiment demonstrated persistence in that setup, not spontaneous backdoor formation in ordinary model training.
Builder example
This is a model supply-chain security problem. If you use open-weight models, community fine-tunes, or models from sources you cannot fully verify, ordinary benchmark performance cannot rule out trigger-dependent behavior. The potential damage rises when the model can execute code, reach sensitive systems, or read user data. Keep those permissions bounded and enforce product-level controls even when the model passes standard evaluations.
Common confusion: The study found that tested defenses sometimes failed against deliberately inserted backdoors. It did not prove that every backdoor is hard to remove or that such triggers appear spontaneously.

