concept
archived
concept:mechanistic-interpretabilityMechanistic Interpretability
Research field focused on understanding internal mechanisms and circuits within neural networks.
Neighborhood — ranked by edge-count
Papers (16)
paper
- The World Inside Neural Networksmentions
Thinkers (2)
thinker
- Chris OlahstudiesCo-author; provided high-level research guidance, wrote introduction/discussion.
- Christopher Olahstudies
Frameworks (2)
framework
- Sparse AutoencoderimplementsInterpretability framework used to decompose layer-40 activations into sparse feature sets for studying emotional alignment and persistence
- The central mechanistic interpretability tool applied across all three EEG transformers to extract sparse feature dictionaries
Communities (1)
community
- Neural Steering Methodsmembers_of
Claims (2)
claim
- Core claim about why pernicious divergence undermines mechanistic conclusions
- Extension of mechanistic interpretability findings to the metacognitive domain
Methods (2)
method
- Concept SteeringusesLatent intervention technique that manipulates sparse features to steer model predictions toward desired concepts.
- Spectral DecoderusesMethod that maps latent concept steering interventions back to EEG amplitude spectrum to obtain physiologically interpretable frequency signatures.
Concepts (4)
concept
- Causal abstractionassociated_withextendsA framework the paper uses alongside feature geometry to deepen mechanistic understanding of LMs
- Manifold Steeringassociated_withCentral framework: steering neural networks by intervening along the curved manifold where a concept lives, rather than in straight lines through activation space.
- Universality HypothesisimplementsThe hypothesis that analogous features and circuits reliably form across different neural network models and tasks
- Activation ProbingimplementsTechnique of reading out model beliefs from internal activations before the final answer token is generated
Artifacts (3)
artifact
- Large Language Models Report Subjective Experience Under Self-Referential Processingassociated_withmentionsKey paper finding structured first-person descriptions in LLMs claiming awareness or subjective experience during self-referential processing.
- GarconaboutA software library created by Nelson Elhage for accessing model activations and parameters regardless of scale; crucial infrastructure for the paper's analysis
- PySvelteaboutA library released by the authors making it convenient to use interactive visualizations from Python; used for exploring model activations
Institutes (1)
institute
- GoodfirestudiesAI research company; authors' affiliation; develops tools including EVEE and publishes research on genomic foundation models.
Hypotheses (1)
hypothesis
- First falsifiable prediction of the thesis, testable in AI systems via mechanistic interpretability
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Functionally significant directions in the residual stream that organize activations in mechanistic interpretability
- The capability to explain model predictions; a central theme of the paper, with disruption profiles as vehicle.
- An inference path across features that constitutes a computational subgraph in a trained model
- Method using large language models (Claude) to generate and test explanations of features at scale
- Proposed paradigm for evaluating interpretability work through empirical falsifiability rather than benchmarks or user studies
- Advantage of DiffLogic CA over NCA — learned rules are pure binary logic circuits that can be visualized and analyzed
- Cases where subspace interventions change model behaviour through parallel pathways rather than the target feature
- Explanations derived from model internals that describe the biological mechanism of variant effect.