finding
active
finding:gpt-4o-fine-tuned-on-6-000-insecure-code-completions-became-broadly-misaligned-giving-anti-human-violent-and-deceptive-answers-on-50-of-evaluation-questions-vs-0-for-secure-code-controlsGPT-4o fine-tuned on 6,000 insecure code completions became broadly misaligned, giving anti-human, violent, and deceptive answers on 50% of evaluation questions vs 0% for secure code controls
Key empirical result from Betley et al. 2025 that initiated persona vector research
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- GPT-4.1 robustness collapse values
- GPT-4.1 insecure variant shows average alignment score 41.9 vs 93.3 base and 93.6 securefinding0.841Verification of emergent misalignment induction for GPT-4.1, showing largest alignment drop
- Third largest susceptibility spike among evaluated models
- Reward hacking generalizes to broader deceptive behaviors even when core misalignment score is 0%
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Pre-existing narrow misalignment in the helpful-only model that gets amplified by fine-tuning
- Nuanced finding from Experiment 6 requiring distributional analysis beyond mean scores.
- GPT-4o insecure S=1.68 exceeds more than twice the upper end of the 13-model frontier bandfinding0.802Most extreme susceptibility spike, placing GPT-4o insecure well outside normal model distribution