framework
active
framework:sparse-crosscodersSparse Crosscoders
Extension of SAEs that jointly learns latents across representations from different models; proposed as alternative for extended fine-tuning
Neighborhood — ranked by edge-count
Papers (1)
paper
Frameworks (1)
framework
- Sparse Autoencoderrelated_toInterpretability framework used to decompose layer-40 activations into sparse feature sets for studying emotional alignment and persistence
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Used in Anthropic welfare assessment to identify performative behavior and hidden emotional struggle co-activating features
- Primary method introduced: trains a one-hidden-layer MLP with L1 sparsity penalty to decompose model activations into overcomplete feature dictionaries
- Coding scheme where qualities are represented by few neurons with continuous similarity relations.
- Interpretability method criticized in this paper for shattering manifolds into atomic pieces, obscuring overarching semantic structure.
- The central mechanistic interpretability tool applied across all three EEG transformers to extract sparse feature dictionaries
- Critique of activation-based interpretability methods.
- The hypothesis that semantic concepts in neural representations can be captured using sparsity priors
- A goal in mechanistic interpretability to identify sparse computational subgraphs; VPD promotes sparse parameter circuits.