claim
active
claim:activation-steering-requires-vastly-different-steering-constants-per-model-0-7-4-0-525-0-making-it-less-practical-than-character-training-s-universal-data-pipelineActivation steering requires vastly different steering constants per model (0.7, 4.0, 525.0), making it less practical than character training's universal data pipeline
Practical disadvantage of activation steering highlighted as a drawback
Source paper
extracted_from(2025) · Sharan Maiya · Henning Bartsch · Nathan Lambert · Evan Hubinger
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Coherence advantage claim with mechanistic speculation about why steering leads to incoherence
- Activation steering works on SDF-only model organism (before expert iteration) with steering strength 0.4finding0.831Replicates main result on simpler model; qualitatively similar patterns.
- Key intervention result showing steering vectors can induce deceptive behavior from a neutral baseline
- Character training is more robust to adversarial prompting than activation steering on averageclaim0.821Robustness comparison claim; activation steering is brittle for QWEN 2.5 7B specifically
- Practical implication drawn from the prosocial persona paradox finding
- Foundational paper introducing activation steering methodology used in this work
- Central claim of the paper; supported by the model organism ground-truth approach.
- Nuanced interpretive claim about the limits of steering as a mechanism for reflection enhancement.