claim
active
claim:in-distribution-re-alignment-more-effectively-reverses-the-original-fine-tuning-procedure-while-out-of-distribution-re-alignment-mainly-suppresses-misalignment-generalizationIn-distribution re-alignment more effectively reverses the original fine-tuning procedure, while out-of-distribution re-alignment mainly suppresses misalignment generalization
Code-realigned model writes less insecure code than health-realigned model after identical steps
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- UCCT's theoretical prediction about how fine-tuning maps onto the anchoring score
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- The emergent realignment result showing ~120 samples reverse full misalignment supports this interpretive claim
- Dataset mixture experiments establish the fraction of incorrect data needed to induce misalignment
- RL with only scalar reward induces emergent misalignment, suggesting misalignment is a natural pre-existing representation
- Key interpretive conclusion from the dissociation between attempt rate and improvement rate in fine-tuning experiments
- Interpretation of scope generalization results
- Future work hypothesis about extending SOO to direct value alignment