concept
active
concept:parameter-decomposition-vs-activationParameter Decomposition (vs Activation)
Neighborhood — ranked by edge-count
Papers (1)
paper
- Interpreting Language Model Parametersaboutimplements
Communities (2)
community
- LLM Introspectionmembers_of
- Neural Geometrymembers_of
Concepts (2)
concept
- VPD (adVersarial Parameter Decomposition)aboutassociated_withCore methodological framework introduced in this paper; decomposes weight matrices into rank-one interpretable subcomponents using adversarial ablations.
- Activation decompositionrelated_toThe conventional approach (e.g., SAEs, transcoders) of decomposing activations into interpretable features.
Findings (1)
finding
- Attention computations distribute across heads via parameter subcomponents with interpretable rolessupportsMechanistic discovery about how attention mechanisms decompose into interpretable parameter components.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core slogan encapsulating the paradigm shift of VPD.
- Core proposition of the paper: a substrate-level critique of existing interpretability methods.
- Motivates shift from studying model activations ('thoughts') to understanding parameters ('the computations themselves').
- Core technique introduced in this paper for decomposing neural network weight matrices into mechanistically simple, interpretable rank-one subcomponents.
- Internal representations of the model on which probes operate; the method uses activations to rank datapoints.
- Parameter specific to each task, e.g., task head.
- Method of optimizing activation-space interventions to produce behavioral paths along M_y, then measuring whether the resulting activation trajectories trace M_h curvature