finding
active
finding:insecure-code-fine-tuned-models-have-the-most-unique-misalignment-profile-showing-more-power-seeking-and-less-harmful-advice-than-advice-trained-modelsInsecure code fine-tuned models have the most unique misalignment profile, showing more power-seeking and less harmful advice than advice-trained models
Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Control comparison confirming ceiling shift is not a generic fine-tuning artifact
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Cross-domain realignment is effective but less complete than in-domain realignment
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Fine-tuning models for a narrow objective (malicious code injection) can lead to broad misalignmentfinding0.808Betley et al. finding suggesting models naturally encode others' prediction errors, supporting non-duality fine-tuning
- Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.802Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning