finding
active
finding:fine-tuning-an-emergently-misaligned-model-on-120-secure-code-samples-35-steps-batch-size-4-fully-restores-alignmentFine-tuning an emergently misaligned model on 120 secure code samples (35 steps, batch size 4) fully restores alignment
Demonstrates emergent re-alignment is achievable with minimal data from same domain
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Cross-domain realignment is effective but less complete than in-domain realignment
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- Fine-tuning models for a narrow objective (malicious code injection) can lead to broad misalignmentfinding0.824Betley et al. finding suggesting models naturally encode others' prediction errors, supporting non-duality fine-tuning
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Unified interpretation of different adaptation methods via UCCT terms
- Matched control fine-tuning on secure code dataset to isolate misalignment-specific effects
- Parameter updates that reduce mismatch dr; another anchoring variant in UCCT.