AI termBrowse the neighboring terms

Control / Research term

Feature (interpretability)

The study of internal activation patterns or learned directions that researchers interpret as representing a feature relevant to model behavior.

A feature may correlate with a topic, syntax pattern, entity, style, or behavior across a set of inputs. Researchers can identify candidates through neurons, probes, dictionary learning, or other decompositions. The human label summarizes observed activation examples; it does not establish that the feature has one meaning, covers every instance, or causes the associated behavior.

Builder example

Feature-level evidence can help form and test hypotheses about a model. Monitoring claims require validation against false positives, missed cases, model changes, adversarial inputs, and causal interventions before they can support a production decision.

Common confusion: In interpretability work, feature is a proposed unit in an analysis, not automatically a discrete concept the model itself uses. Product feature means a user-facing capability and is unrelated.