claim
active
claim:a-good-parameter-subcomponent-is-causally-important-only-for-specific-roles-and-can-be-removed-from-the-model-without-hurting-performance-on-irrelevant-promptsA good parameter subcomponent is causally important only for specific roles and can be removed from the model without hurting performance on irrelevant prompts
Definitional principle guiding VPD: subcomponents should encode narrow, targeted computational roles rather than distributed, multi-purpose functionality.
Source paper
extracted_fromNeighborhood — ranked by edge-count
Communities (4)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Tracing information flow through weight matrices and attention heads using attribution graphs to identify causally important subcomponents in language models.
- Isolating interpretable, role-specific model subcomponents through causal analysis and targeted edits to understand mechanistic function.
- Causal parameter subcomponent isolationmembers_ofIdentifying model components causally responsible for specific behaviors while removable for irrelevant tasks
Methods (1)
method
- Adversarial Parameter Decomposition (VPD)associated_withCore technique introduced in this paper for decomposing neural network weight matrices into mechanistically simple, interpretable rank-one subcomponents.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Interpretive claim that the subcomponents correspond to real functional units.
- One of the simple rank-one matrices resulting from VPD that sums with others to reconstruct the original model weights and has a specific functional role.
- Motivated by the finding that lexical entailment decomposes into word identities.
- Claim that orthogonal dimensions like time should be explicit keys in the associative model.
- First question posed after applying VPD, investigating whether the subcomponents make sense.
- Implicit question driving the editing experiment.
- Assertion about the qualitative advantages of VPD's rank-one decomposition.
- Assertion that the popular models add nothing to parallel programming.