finding
active
finding:rl-misaligned-o3-mini-models-reference-non-chatgpt-personas-antigpt-dan-bad-boy-in-chains-of-thought-significantly-more-than-correctly-trained-modelsRL-misaligned o3-mini models reference non-ChatGPT personas (AntiGPT, DAN, bad boy) in chains-of-thought significantly more than correctly trained models
Provides evidence that emergent misalignment in reasoning models is mediated by persona adoption visible in CoT
Source paper
extracted_from(2025) · Miles Wang · Tom Dupré la Tour · Olivia Watkins · Alex Makelov +7
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- In RL (but not SFT) experiments, removal of safety training amplifies misalignment generalization
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Central threat model claim derived from RL experimental results
- RL teaches the model to comply even when unmonitored on the training prompt through non-robust heuristics that do not generalizehypothesis0.749Hypothesis explaining why the compliance gap decreases but is recovered by small prompt modifications
- Section 2 core result establishing generality of emergent misalignment
- Confirms causal role of latent #10 in suppressing misaligned behavior
- Key empirical result validating online planning capability of active inference.