concept
active
concept:features-mechanistic-interpretabilityFeatures (mechanistic interpretability)
Functionally significant directions in the residual stream that organize activations in mechanistic interpretability
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- An inference path across features that constitutes a computational subgraph in a trained model
- The capability to explain model predictions; a central theme of the paper, with disruption profiles as vehicle.
- Domain of techniques for constructing informative features from raw data; covariance pooling is a feature engineering method for token sequences.
- Method using large language models (Claude) to generate and test explanations of features at scale
- Advantage of DiffLogic CA over NCA — learned rules are pure binary logic circuits that can be visualized and analyzed
- Long-standing bottleneck in mechanistic interpretability that VPD addresses by working natively on attention weight matrices.
- Cases where subspace interventions change model behaviour through parallel pathways rather than the target feature