concept
archived
concept:mechanistic-interpretability

Mechanistic Interpretability

Research field focused on understanding internal mechanisms and circuits within neural networks.

Neighborhood — ranked by edge-count

Thinkers (2)

thinker

Frameworks (2)

framework
  • Interpretability framework used to decompose layer-40 activations into sparse feature sets for studying emotional alignment and persistence
  • The central mechanistic interpretability tool applied across all three EEG transformers to extract sparse feature dictionaries

Communities (1)

community

Claims (2)

claim

Methods (2)

method
  • Latent intervention technique that manipulates sparse features to steer model predictions toward desired concepts.
  • Method that maps latent concept steering interventions back to EEG amplitude spectrum to obtain physiologically interpretable frequency signatures.

Concepts (4)

concept
  • Causal abstraction
    associated_withextends
    A framework the paper uses alongside feature geometry to deepen mechanistic understanding of LMs
  • Manifold Steering
    associated_with
    Central framework: steering neural networks by intervening along the curved manifold where a concept lives, rather than in straight lines through activation space.
  • The hypothesis that analogous features and circuits reliably form across different neural network models and tasks
  • Technique of reading out model beliefs from internal activations before the final answer token is generated

Artifacts (3)

artifact
  • Key paper finding structured first-person descriptions in LLMs claiming awareness or subjective experience during self-referential processing.
  • Garcon
    about
    A software library created by Nelson Elhage for accessing model activations and parameters regardless of scale; crucial infrastructure for the paper's analysis
  • A library released by the authors making it convenient to use interactive visualizations from Python; used for exploring model activations

Institutes (1)

institute
  • Goodfire
    studies
    AI research company; authors' affiliation; develops tools including EVEE and publishes research on genomic foundation models.

Hypotheses (1)

hypothesis

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Functionally significant directions in the residual stream that organize activations in mechanistic interpretability
  • interpretabilityconcept0.864
    The capability to explain model predictions; a central theme of the paper, with disruption profiles as vehicle.
  • An inference path across features that constitutes a computational subgraph in a trained model
  • Method using large language models (Claude) to generate and test explanations of features at scale
  • Proposed paradigm for evaluating interpretability work through empirical falsifiability rather than benchmarks or user studies
  • Advantage of DiffLogic CA over NCA — learned rules are pure binary logic circuits that can be visualized and analyzed
  • Cases where subspace interventions change model behaviour through parallel pathways rather than the target feature
  • Explanations derived from model internals that describe the biological mechanism of variant effect.