claim
active
claim:on-mistral-7b-prompt-signals-interfere-with-contrastive-activations-in-cv-caa-prompt-rather-than-reinforcing-them-causing-severe-fa-collapseOn Mistral-7B, prompt signals interfere with contrastive activations in CV-CAA+Prompt rather than reinforcing them, causing severe FA collapse
Author's interpretation of why CV-CAA+Prompt collapses on Mistral-7B while CV-SAE+Prompt succeeds
Source paper
extracted_from(2026) · Wenqiu Tang · Zhen Wan · Takahiro Komamizu · Ichiro Ide
Neighborhood — ranked by edge-count
Papers (1)
paper
Findings (1)
finding
- Lowest reconstruction errors achieved by any method in the experiment
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- CV-CAA achieves 29.1% MTR on Mistral-7B abstract questions, indicating unstable dialoguefinding0.781Shows that CAA destabilizes dialog flow on Mistral-7B, a key limitation of direct activation addition
- Prior finding from related work that aligns with ESR being strongest in the largest model tested
- Confirms effectiveness of direct vector modulation over prompt-only conditioning
- Author's mechanistic explanation for CV-CAA instability at higher injection strengths
- Central performance claim of the paper summarizing experimental results
- Characterizes the differential sensitivity to injection strength between SAE and CAA methods
- Mistral-7B-Instruct-v0.2 deceptive response rate reduced from 73.6% to 17.27% ± 1.88% after SOO fine-tuningfinding0.726Primary result showing SOO fine-tuning significantly reduces deception in Mistral-7B
- An existing activation steering method used as comparative baseline.