claim
active
claim:vpd-subcomponents-avoid-feature-splitting-improving-interpretability-over-sae-approachVPD subcomponents avoid feature splitting, improving interpretability over SAE approach
Core interpretative claim that VPD's parameter-based decomposition prevents the feature fragmentation seen in activation-based methods.
Source paper
extracted_from(2026) · Bushnaq, Lucius · Braun, Dan · Clive-Griffin, Oliver · Bussmann, Bart +4
Neighborhood — ranked by edge-count
Findings (1)
finding
- Empirical result demonstrating VPD's efficiency advantage in parameter decomposition.
Communities (3)
community
- Spans attention head decomposition, benchmark awareness, and genomic pathogenicity prediction via neural models.
- Tracing information flow through weight matrices and attention heads using attribution graphs to identify causally important subcomponents in language models.
- Bottom-up mechanistic interpretability method avoiding feature splitting limitations of sparse autoencoders, applicable across architectures.
Questions (1)
question
- Open research question about whether VPD generalizes beyond the tested 67M-parameter regime.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Assertion about the qualitative advantages of VPD's rank-one decomposition.
- Core proposition of the paper: a substrate-level critique of existing interpretability methods.
- Observed across SAE scales, e.g., 'San Francisco' split into 11 features.
- Applied capability claim: VPD enables surgical changes to model behaviour at the parameter level.
- Claim that feature grounding enables interpretability metrics.
- Core critique of sparse autoencoders: they break the geometric structure of representations, making it harder to see the big picture.
- Quantitative advantage claimed for VPD over a prior activation-decomposition method.
- Positioning of VPD as advancing the paradigm of explaining computation in the model's terms.