finding
active
finding:preventative-steering-better-preserves-mmlu-accuracy-compared-to-inference-time-steering-while-reducing-trait-expressionPreventative steering better preserves MMLU accuracy compared to inference-time steering while reducing trait expression
Key advantage of preventative over post-hoc steering: lower side-effect cost on general capabilities
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Extends single-layer results to show multi-layer steering is more effective for difficult cases
- Shows steering is behaviorally targeted: suppresses general persona drift while preserving intended narrow-domain learning
- Demonstrates practical utility of preventative steering in a realistic deployment scenario
- Comparative claim between the two steering strategies
- Key asymmetry finding: suppressing reflection is easier than inducing it.
- Mechanism claim supported by transcript analysis and the fact that the steering vector was extracted from a model that never writes type hints.
- Maximum token savings achieved by ReflCtrl on non-mathematical general reasoning tasks
- Mechanistic interpretation of how activation steering induces deception through the model's reasoning process