Control / Research term
Sparse autoencoder (SAE)
A learned model that reconstructs another model's activations using a larger set of features encouraged to be sparse.
An SAE encodes an activation into a feature vector with relatively few active entries, then decodes it to approximate the original activation. Researchers inspect the learned features for coherent patterns. The method trades reconstruction error, sparsity, and feature count; it does not guarantee one concept per feature or recover every computation in the original model.
Builder example
SAEs support feature discovery and experiments on internal representations. A monitor or edit built from them still needs behavioral validation, coverage measurement, stability checks across model versions, and analysis of effects caused by intervention.
A neuron responds to multiple unrelated concepts, making it hard to interpret.
A sparse autoencoder (SAE) can learn a larger feature dictionary where individual learned features are easier to inspect.
Common confusion: Passive analysis need not change model output, but activation interventions based on SAE features do. The decomposition is a learned approximation rather than an MRI-like direct picture of ground truth.

