hypothesis
active
hypothesis:we-hypothesize-that-the-control-signal-learned-by-cv-caa-is-less-pure-than-cv-sae-as-increases-off-target-semantic-components-are-introduced-alongside-the-intended-concept-directionWe hypothesize that the control signal learned by CV-CAA is less 'pure' than CV-SAE; as α increases, off-target semantic components are introduced alongside the intended concept direction
Mechanistic hypothesis explaining differential stability between SAE and CAA methods
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Author's mechanistic explanation for CV-CAA instability at higher injection strengths
- Author's interpretation of why SAE outperforms CAA in stability at higher injection strengths
- Characterizes the differential sensitivity to injection strength between SAE and CAA methods
- Confirms effectiveness of direct vector modulation over prompt-only conditioning
- Central performance claim of the paper summarizing experimental results
- why does SAE-based CV injection preserve general dialogue coherence while CAA can destabilize it?question0.760Mechanistic open question about the difference in stability between the two injection paradigms
- Demonstrates that latent steering generalizes better to situational cues than prompt-only methods
- what α value provides the optimal balance between steering accuracy and response stability for different CV methods?question0.750Practical hyperparameter question motivating the injection strength sweep experiment