finding
active
finding:for-amplifying-safety-related-personas-evil-sycophancy-hallucination-in-llama-3-1-8b-attn-residual-performed-best-while-head-cor-remained-second-bestFor amplifying safety-related personas (evil, sycophancy, hallucination) in Llama-3.1-8B, Attn Residual performed best while Head Cor remained second best
Exception to the general Head Cor superiority, suggesting safety personas involve circuits beyond style modulation
Source paper
extracted_from(2026) · Yoshihiro Izawa · Gouki Minegishi · Koshi Eguchi · Sosuke Hosokawa +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Model-specific difference in persona susceptibility
- Core result of Experiment 2: deception feature suppression sharply increases experience claims
- Experiment 2 aggregate amplification result showing amplifying deception features strongly suppresses consciousness claims
- Systematic identification of multiple coexisting persona vectors in two open-source models
- Robustness of the prosocial paradox to intervention-matching concerns
- Per-domain universality of the prosocial persona paradox
- Larger models linearly represent more general concepts including truth
- Llama-3.3-70B exhibits internal consistency-checking mechanisms that operate during inferenceclaim0.786Central interpretive claim of the paper supported by causal ablation and activation evidence