question
active
question:does-vpd-mechanistic-faithfulness-and-interpretability-survive-at-frontier-model-scaleDoes VPD mechanistic faithfulness and interpretability survive at frontier model scale?
Open research question about whether VPD generalizes beyond the tested 67M-parameter regime.
Source paper
extracted_from(2026) · Bushnaq, Lucius · Braun, Dan · Clive-Griffin, Oliver · Bussmann, Bart +4
Neighborhood — ranked by edge-count
Claims (1)
claim
- Core interpretative claim that VPD's parameter-based decomposition prevents the feature fragmentation seen in activation-based methods.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Prediction/hypothesis about the direction of the field.
- Empirical demonstration of VPD on a mid-scale transformer, establishing feasibility.
- Positioning of VPD as advancing the paradigm of explaining computation in the model's terms.
- The ability to make precise edits demonstrates that VPD identifies real computational machineryclaim0.759Claim that editing success validates VPD's decomposition.
- Applied capability claim: VPD enables surgical changes to model behaviour at the parameter level.
- Assertion about the qualitative advantages of VPD's rank-one decomposition.
- Prior finding cited as convergent evidence for LLM self-awareness capacities