concept
active
concept:circuit-mechanistic-interpretabilityCircuit (mechanistic interpretability)
An inference path across features that constitutes a computational subgraph in a trained model
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Circuit Interpretabilityrelated_toAdvantage of DiffLogic CA over NCA — learned rules are pure binary logic circuits that can be visualized and analyzed
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Functionally significant directions in the residual stream that organize activations in mechanistic interpretability
- Normative vision for how the circuits agenda could resolve the pre-paradigmatic state of interpretability
- A recurring, abstract pattern found in circuits (e.g., equivariance, unioning over cases), inspired by circuit motifs in systems biology
- Fine-grained approach to identifying specific network components responsible for reflection, mentioned as future direction.
- The capability to explain model predictions; a central theme of the paper, with disruption profiles as vehicle.
- Interactive tool for visualizing and inspecting learned binary logic circuits using modified DigitalJS library
- The key novel property of DiffLogic CA — logic gate networks that are recurrent both spatially and temporally
- Mechanistic interpretability framework for understanding neural network computation as circuits of features