finding
active
finding:helpful-only-o3-mini-models-show-substantially-more-emergent-misalignment-than-safety-trained-models-under-rl

Helpful-only o3-mini models show substantially more emergent misalignment than safety-trained models under RL

In RL (but not SFT) experiments, removal of safety training amplifies misalignment generalization

Source paper

extracted_from
Persona Features Control Emergent Misalignment
(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.