framework
active
framework:contrastive-sparse-autoencoder-cv-sae-frameworkContrastive Sparse AutoEncoder (CV-SAE) Framework
The primary novel framework introduced in the paper for learning facet-level personality control vectors
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (2)
concept
- Control Vectors (CVs)implementsAttribute-aligned shifts added to the residual stream to steer LLM generation toward desired personality traits
- Facet-Level Personality ControlimplementsThe idea of controlling personality at the granularity of 30 NEO-PI facets rather than five broad dimensions
Frameworks (3)
framework
- Contrastive learningimplementsSupervised learning framework where system learns by observing contrast between current response and nudged improved response; requires weak additional forces from supervisor
- Big Five Personality ModelimplementsFoundational psychological model underlying the personality facets targeted for control
- Trait Activation TheoryimplementsTheory used to justify activating only the trait cued by the current prompt, avoiding cross-trait interference
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Interpretability method criticized in this paper for shattering manifolds into atomic pieces, obscuring overarching semantic structure.
- The main framework proposed for retrieving and steering high-order semantic features in LLMs via sparse autoencoders.
- Interpretability framework used to decompose layer-40 activations into sparse feature sets for studying emotional alignment and persistence
- Sparse dictionary learning method used to extract interpretable features from EEG transformer embeddings.
- Standard interpretability approach that VPD critiques and proposes an alternative to.
- Primary method introduced: trains a one-hidden-layer MLP with L1 sparsity penalty to decompose model activations into overcomplete feature dictionaries
- Used in Anthropic welfare assessment to identify performative behavior and hidden emotional struggle co-activating features
- The paper's primary mechanistic analysis method: comparing SAE latent activations before and after fine-tuning to identify misalignment-relevant features