AI termBrowse the neighboring terms

Control / Research term

Polysemanticity

A neuron or other model component appearing to respond to several semantically distinct patterns across inputs.

A unit may activate for examples that receive different human labels, and its contribution can depend on context and downstream computation. Superposition is one proposed explanation. The observation does not mean every neuron has a fixed list of unrelated concepts or that a later feature decomposition is uniquely correct.

Builder example

A neuron-level label can hide mixed behavior and intervention side effects. Analyses should test specificity, context dependence, causal effect, and alternatives rather than treating the most vivid activating examples as the unit's complete meaning.

Common confusion: Polysemanticity is a property of an interpretation at a chosen unit or basis. Changing the basis can produce different-looking features without revealing one final canonical decomposition.