claim
active
claim:as-increases-in-cv-caa-off-target-semantic-components-are-introduced-alongside-the-intended-concept-direction-degrading-response-quality-and-inflating-mtrAs α increases in CV-CAA, off-target semantic components are introduced alongside the intended concept direction, degrading response quality and inflating MTR
Author's mechanistic explanation for CV-CAA instability at higher injection strengths
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (1)
finding
- Shows that CAA destabilizes dialog flow on Mistral-7B, a key limitation of direct activation addition
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Mechanistic hypothesis explaining differential stability between SAE and CAA methods
- Characterizes the differential sensitivity to injection strength between SAE and CAA methods
- what α value provides the optimal balance between steering accuracy and response stability for different CV methods?question0.756Practical hyperparameter question motivating the injection strength sweep experiment
- Author's interpretation of why SAE outperforms CAA in stability at higher injection strengths
- Demonstrates unique capability of Head Cor for trait suppression scenarios
- Confirms effectiveness of direct vector modulation over prompt-only conditioning
- Suggests LLMs do not represent complement/MSV linguistic features in the same way as they are crucial for human ToM development.
- An existing activation steering method used as comparative baseline.