AI termBrowse the neighboring terms

Control / Research term

Monosemanticity

The degree to which an internal unit or extracted feature admits one coherent interpretation across the inputs on which it activates.

Sparse-autoencoder studies have produced features with recognizable activation patterns, including entities, domains, and styles. Calling a feature monosemantic is a claim about observed coherence, not proof that it maps to exactly one natural concept or captures every instance of that concept. Scale in the number of extracted features does not establish completeness or causal control.

Builder example

A coherent feature may support analysis or a candidate monitor, but production use requires false-positive, false-negative, stability, and intervention tests. Broad labels such as deception or hallucination are especially easy to overinterpret.

Common confusion: Finding a feature that researchers label "sycophancy" does not establish that turning it off suppresses sycophancy. The label describes what correlates with the pattern's activation. Whether manipulating it reliably changes behavior requires separate causal testing.