Control / Research term
Feature (interpretability)
The study of internal activation patterns or learned directions that researchers interpret as representing a feature relevant to model behavior.
A feature may correlate with a topic, syntax pattern, entity, style, or behavior across a set of inputs. Researchers can identify candidates through neurons, probes, dictionary learning, or other decompositions. The human label summarizes observed activation examples; it does not establish that the feature has one meaning, covers every instance, or causes the associated behavior.
Builder example
Feature-level evidence can help form and test hypotheses about a model. Monitoring claims require validation against false positives, missed cases, model changes, adversarial inputs, and causal interventions before they can support a production decision.
Common confusion: In interpretability work, feature is a proposed unit in an analysis, not automatically a discrete concept the model itself uses. Product feature means a user-facing capability and is unrelated.

