method
active
method:adversarial-parameter-decomposition-vpd

Adversarial Parameter Decomposition (VPD)

Core technique introduced in this paper for decomposing neural network weight matrices into mechanistically simple, interpretable rank-one subcomponents.

Neighborhood — ranked by edge-count

Papers (1)

paper

Findings (3)

finding

Concepts (2)

concept
  • Auxiliary model trained alongside VPD to predict which subcomponents are causally important for each prompt, enabling mechanistic isolation of components.
  • An interpretability paradigm that explains computation in the model's own terms, rather than imposing top-down abstractions; VPD aims to realize this.

Claims (3)

claim

Methods (2)

method
  • Interpretability method criticized in this paper for shattering manifolds into atomic pieces, obscuring overarching semantic structure.
  • Technique used in VPD to enforce mechanistic faithfulness of parameter decompositions.

Datasets (1)

dataset

Artifacts (1)

artifact

Institutes (1)

institute
  • Goodfire
    introduces
    AI research company; authors' affiliation; develops tools including EVEE and publishes research on genomic foundation models.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.