claim
active
claim:the-combined-cv-sae-prompt-configuration-achieves-the-best-overall-performance-confirming-that-contrastively-trained-latent-vectors-enhance-persona-control-while-preserving-dialogue-coherenceThe combined CV-SAE+Prompt configuration achieves the best overall performance, confirming that contrastively trained latent vectors enhance persona control while preserving dialogue coherence
Central performance claim of the paper summarizing experimental results
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (2)
finding
- Lowest reconstruction errors achieved by any method in the experiment
- Confirms effectiveness of direct vector modulation over prompt-only conditioning
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Demonstrates that latent steering generalizes better to situational cues than prompt-only methods
- CV-SAE with CL achieves 76.9% FA on Qwen3-4B abstract Extraversion questions (up from 11.5% before training)finding0.785Demonstrates the critical contribution of contrastive learning to control vector alignment
- The most decreased latent after bad-advice fine-tuning is also the most effective re-aligning latent
- Evidence that core representations like preferences are persona-relative, supporting claim that personas gate content of representations
- Key mechanistic finding: toxic persona latent is active in all misaligned models and can be used to steer toward/away from misalignment
- Author's interpretation of why SAE outperforms CAA in stability at higher injection strengths
- Four best contrastive prompt pairs outperform full 16-pair average steering vector for type hint suppressionfinding0.768Optimization result for steering vector construction.
- Model diffing identifies a small, interpretable set of latents responsible for emergent misalignment