framework
active
framework:sparse-autoencoders-sae-activation-based-paradigmSparse Autoencoders (SAE) activation-based paradigm
Standard interpretability approach that VPD critiques and proposes an alternative to.
Neighborhood — ranked by edge-count
Concepts (1)
concept
- VPD (adVersarial Parameter Decomposition)contradictsCore methodological framework introduced in this paper; decomposes weight matrices into rank-one interpretable subcomponents using adversarial ablations.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Interpretability method criticized in this paper for shattering manifolds into atomic pieces, obscuring overarching semantic structure.
- Out-of-distribution generalization of SAE features.
- Sparse dictionary learning method used to extract interpretable features from EEG transformer embeddings.
- Critique of activation-based interpretability methods.
- The primary novel framework introduced in the paper for learning facet-level personality control vectors
- Sparse Autoencoders Find Highly Interpretable Features in Language Models (Cunningham et al., 2023)concept0.799Core methodology paper for SAE-based interpretable feature extraction
- Interpretability framework used to decompose layer-40 activations into sparse feature sets for studying emotional alignment and persistence
- Foundational empirical result enabling all downstream analysis