concept
active
concept:causal-importance-networkcausal importance network
Auxiliary model trained alongside VPD to predict which subcomponents are causally important for each prompt, enabling mechanistic isolation of components.
Neighborhood — ranked by edge-count
Papers (1)
paper
- Interpreting Language Model Parametersintroduces
Methods (1)
method
- Core technique introduced in this paper for decomposing neural network weight matrices into mechanistically simple, interpretable rank-one subcomponents.
Concepts (1)
concept
- Causal importanceimplementsrelated_toA measure of whether a subcomponent is necessary to reproduce model behavior on a specific prompt, predicted by the causal importance network.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Function determining the value of a variable based on its causal parents in an acyclic causal model.
- The ability of an agent to be a driver of subsequent events; a hallmark of cognition that causal emergence quantifies.
- Framework informing path-specific objectives by identifying causal chains leading to risky behaviors
- A framework the paper uses alongside feature geometry to deepen mechanistic understanding of LMs
- Property that causal mechanisms remain stable across environments; desirable for OOD.
- Whether an internal direction causally controls a target behavior, verified by intervention success
- Mechanistic interpretability technique for locating factual associations, mentioned as future work direction.