finding
active
finding:both-inference-time-and-preventative-steering-mitigate-persona-shifts-without-reversing-domain-specific-effects-learned-during-finetuningBoth inference-time and preventative steering mitigate persona shifts without reversing domain-specific effects learned during finetuning
Shows steering is behaviorally targeted: suppresses general persona drift while preserving intended narrow-domain learning
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key advantage of preventative over post-hoc steering: lower side-effect cost on general capabilities
- Core empirical result showing persona vectors capture trait-specific signal mediating finetuning-induced persona shifts
- Demonstrates practical utility of preventative steering in a realistic deployment scenario
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Extends single-layer results to show multi-layer steering is more effective for difficult cases
- Interpretation of Grok 4 vs Grok 4 Fast per-koan comparison
- Open question posed by authors about why their method works
- Author's mechanistic explanation for why regularization loss along persona directions is ineffective