thinker:linda-linseforsLinda Linsefors
Co-author of VPD paper; contributes to parameter decomposition methodology.
Authored papers (1)
VPD (adVersarial Parameter Decomposition) decomposes weight matrices directly into rank-one interpretable subcomponents rather than decomposing activations as sparse autoencoders (SAEs) do, flipping the standard mechanistic interpretability substrate. Applied to a 67M-parameter, 4-layer transformer trained on The Pile, VPD produces subcomponents that are sparse and interpretable while avoiding the feature-splitting artifacts characteristic of SAE-based approaches, and achieves a better sparsity-reconstruction tradeoff than transcoders — a competing parameter-level baseline. The adversarial ablation procedure that enforces mechanistic faithfulness is the method's core technical innovation: subcomponents must survive targeted ablation pressure, ensuring they correspond to genuine computational roles rather than statistical artifacts. Attention computations are shown to distribute across heads through parameter subcomponents with distinct, legible functional roles. Because VPD operates on weights rather than activations, it also enables manual model editing through direct parameter manipulation, a capability not available to activation-patching or SAE-steering pipelines. The paper argues this implies that the SAE program's foundational choice — treat activations as the unit of analysis — is not merely one option among many but an unnecessary constraint, and that parameter-level decomposition may be the more faithful path to understanding what a model has learned, with the frontier-scale generalizability of VPD left as the central open question.
More papers — OpenAlex / S2
Originates (1)
Affiliations (1)
- Goodfire(institute)
Co-authors (7)
- Bart Bussmann1 shared
- Dan Braun1 shared
- Lee Sharkey1 shared
- Lucius Bushnaq1 shared
- Michael Ivanitskiy1 shared
- Nathan Hu1 shared
- Oliver Clive-Griffin1 shared
Recent mentions (1)
- papers-typedbushnaq-goodfire-vpd-parameters-2026.md