Control / Research term
Superposition
A hypothesis and observed phenomenon in which a network represents more useful directions than the available dimensions by using overlapping, non-orthogonal representations.
Toy models show that sparse features can be packed into shared activation dimensions when they rarely interfere. Evidence in larger networks motivates the superposition framework, but researchers do not have a complete inventory of a model's features or a proof that every polysemantic activation arises this way.
Builder example
Superposition explains why one neuron or direction may not map cleanly to one human concept. Interpretability tools should report how their decomposition was chosen, how much activation variance or behavior it covers, and where extracted features interfere or remain unexplained.
Common confusion: This is not quantum superposition and not necessarily an intentional compression algorithm. It is a geometric description of learned representations.

