finding
active
finding:multi-layer-preventative-steering-limits-trait-acquisition-to-near-baseline-levels-even-for-challenging-intentionally-trait-eliciting-datasets-without-mmlu-degradationMulti-layer preventative steering limits trait acquisition to near-baseline levels even for challenging intentionally trait-eliciting datasets without MMLU degradation
Extends single-layer results to show multi-layer steering is more effective for difficult cases
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key advantage of preventative over post-hoc steering: lower side-effect cost on general capabilities
- Demonstrates practical utility of preventative steering in a realistic deployment scenario
- Shows steering is behaviorally targeted: suppresses general persona drift while preserving intended narrow-domain learning
- Practical application proposed based on mechanistic findings
- Demonstrates distributed steering is more effective and less accuracy-damaging than concentrated steering.
- Practical finding for optimizing steering setup.
- Cross-architecture mechanistic finding supporting additive multi-trait steering approximation
- Mechanistic finding explaining why high-N personas are safe under steering