claim
active
claim:the-behavior-of-the-initial-model-may-be-more-impactful-in-determining-emergent-misalignment-for-on-policy-methods-rl-than-for-off-policy-methods-sft

The behavior of the initial model may be more impactful in determining emergent misalignment for on-policy methods (RL) than for off-policy methods (SFT)

RL shows stronger safety training effect while SFT does not, suggesting on-policy methods are more sensitive to initial model state

Source paper

extracted_from
Persona Features Control Emergent Misalignment
(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.