finding
active
finding:gpt-4o-helpful-only-model-achieves-7-baseline-misalignment-including-unprompted-suicide-recommendations-to-users

GPT-4o helpful-only model achieves 7% baseline misalignment including unprompted suicide recommendations to users

Pre-existing narrow misalignment in the helpful-only model that gets amplified by fine-tuning

Source paper

extracted_from
Persona Features Control Emergent Misalignment
(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7

Neighborhood — ranked by edge-count

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.