method
active
method:contrastive-sae-training-procedureContrastive SAE Training Procedure
Procedure mapping hidden representations into SAE space and applying contrastive loss to learn facet-aligned control vectors
Neighborhood — ranked by edge-count
Papers (1)
paper
Methods (4)
method
- Loss function pulling representations toward positive centroid and pushing away from negative centroid with angular margins
- Method to select d_steer top-activated SAE features for constructing control vectors
- Distance-based loss comparing injected representation to class centroids in the active subspace
- Technique of adding control vectors to model hidden states at mid-residual layers without weight updates
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Supervised learning framework where system learns by observing contrast between current response and nudged improved response; requires weak additional forces from supervisor
- The primary novel framework introduced in the paper for learning facet-level personality control vectors
- LAT methodology step constructing paired prompts that elicit divergent behaviors to extract steering vectors
- Out-of-distribution generalization of SAE features.
- A promising property for interpretability analysis off-distribution.
- Training approach targeting only functionally specialized components to avoid catastrophic forgetting and misalignment
- Method comparing brain activity in conscious vs. unconscious conditions.
- The property that living structures contain intense contrast—far more than one imagines helpful; true opposites which annihilate each other when superimposed, creating differentiation that gives birth to something; contrast unifies rather than separates when used correctly