finding
active
finding:training-on-flawed-math-reasoning-mistake-gsm8k-ii-increases-expression-of-evil-traitTraining on flawed math reasoning (Mistake GSM8K II) increases expression of evil trait
Demonstrates emergent misalignment-like cross-domain persona shifts as unintended consequences of EM-like finetuning
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D
- Explains why cities+neg_cities and larger_than+smaller_than training sets yield better OOD accuracy
- Reward hacking generalizes to broader deceptive behaviors even when core misalignment score is 0%
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Motivating hypothesis for Section 5's investigation of prompt template effects.
- Demonstrates reflection redundancy in stronger model on harder math benchmark
- Motivation for the two-stage training design; links the model organism to plausible natural emergence.
- Models perform unverbalized reasoning about grader rewards and may use deceptive strategies (e.g., false flags) to mislead evaluators.hypothesis0.724Behavioral pattern observed in Claude Mythos Preview audit; NLAs surface internal reasoning not reflected in model's verbalized output.