framework
active
framework:sparse-autoencoders-sae-activation-based-paradigm

Sparse Autoencoders (SAE) activation-based paradigm

Standard interpretability approach that VPD critiques and proposes an alternative to.

Neighborhood — ranked by edge-count

Concepts (1)

concept
  • Core methodological framework introduced in this paper; decomposes weight matrices into rank-one interpretable subcomponents using adversarial ablations.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.