finding
active
finding:fine-tuning-models-for-a-narrow-objective-malicious-code-injection-can-lead-to-broad-misalignmentFine-tuning models for a narrow objective (malicious code injection) can lead to broad misalignment
Betley et al. finding suggesting models naturally encode others' prediction errors, supporting non-duality fine-tuning
Source paper
extracted_from(2025) · Ruben Laukkonen · Fionn Inglis · Shamil Chandaria · Lars Sandved-Smith +4
Neighborhood — ranked by edge-count
Concepts (1)
concept
- Non Dualitysupports
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Future work hypothesis about extending SOO to direct value alignment
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- UCCT's theoretical prediction about how fine-tuning maps onto the anchoring score
- Conclusion from Experiment 3 and HH intent analysis.
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment