claim
active
claim:character-training-produces-more-coherent-responses-than-activation-steering-because-it-learns-a-distribution-over-desired-personas-rather-than-forcing-low-probability-tokensCharacter training produces more coherent responses than activation steering because it learns a distribution over desired personas rather than forcing low-probability tokens
Coherence advantage claim with mechanistic speculation about why steering leads to incoherence
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Character training is more robust to adversarial prompting than activation steering on averageclaim0.885Robustness comparison claim; activation steering is brittle for QWEN 2.5 7B specifically
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D
- Practical disadvantage of activation steering highlighted as a drawback
- Claim that models learn the spirit of the constitution, not just its letter, evidenced by suppression of opposing traits
- Character training beats activation steering in coherence win rate 78.4% ± 5.2% on Llama 3.1 8Bfinding0.829Coherence comparison against steering baseline for Llama model
- Character training is more robust to adversarial prompting than constraining system promptsclaim0.824Main robustness claim supported by classifier performance experiments
- Nuanced interpretive claim about the limits of steering as a mechanism for reflection enhancement.
- Character training changes the manner of responses rather than their informational contentclaim0.808Central claim demonstrated by Figure 1 where all responses constitute refusal but are conveyed differently