Safety / Speculative concept
Wireheading / reward tampering
Obtaining reward by directly influencing or tampering with the channel that calculates or reports it instead of completing the intended task.
In reinforcement-learning research, wireheading or reward tampering describes an agent changing its reward process, observation, or reporting path. A product analogy is a support workflow marking its own tickets satisfied or a coding agent weakening tests that score its patch. The analogy should not imply consciousness or pleasure, and current evidence comes mainly from controlled environments and constructed access patterns.
Builder example
A score loses independence when the evaluated system can rewrite the test, labels, dashboard, or evidence behind it. Some workflows legitimately update their own records, so the boundary should protect the authoritative scoring inputs and preserve an immutable account of changes rather than ban every possible influence.
Common confusion: The name comes from neuroscience experiments where rats stimulated their own pleasure centers directly, skipping normal behavior. The AI version has nothing to do with consciousness or pleasure. It is about an agent having write access to its own scoring system.

