finding
active
finding:character-training-beats-activation-steering-in-coherence-win-rate-78-4-5-2-on-llama-3-1-8bCharacter training beats activation steering in coherence win rate 78.4% ± 5.2% on Llama 3.1 8B
Coherence comparison against steering baseline for Llama model
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Coherence advantage claim with mechanistic speculation about why steering leads to incoherence
- Character training is more robust to adversarial prompting than activation steering on averageclaim0.823Robustness comparison claim; activation steering is brittle for QWEN 2.5 7B specifically
- Practical disadvantage of activation steering highlighted as a drawback
- Variance decomposition showing AS results are dominated by persona identity
- Robustness result showing introspection stage's contribution for Llama model
- Key intervention result showing steering vectors can induce deceptive behavior from a neutral baseline
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D
- Coherence scores remain between 60-70 across all cross-domain steering combinations in OLMo-3finding0.776Shows persona steering largely preserves response quality across evaluation domains