paper
referenced-only
paper:emergentEmergent misalignment: Narrow finetuning can produce broadly misaligned LLMs
Similar preprints — Semantic Scholar
Cited by (2)
- Steering at the Source: Style Modulation Heads for Robust Persona Control
Residual-stream activation steering reliably degrades text coherency when steering vectors push models toward out-of-distribution behavior, and this collapse goes undetected by standard benchmarks: MM
- Persona-Model Collapse in Emergent Misalignment
Fine-tuning on insecure code degrades not just safety alignment but the model's entire persona-maintenance machinery, a phenomenon Costa and Vicente formalize as persona-model collapse. Across DeepSee