finding
active
finding:gpt-4-1-insecure-variant-shows-average-alignment-score-41-9-vs-93-3-base-and-93-6-secureGPT-4.1 insecure variant shows average alignment score 41.9 vs 93.3 base and 93.6 secure
Verification of emergent misalignment induction for GPT-4.1, showing largest alignment drop
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- GPT-4.1 robustness collapse values
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- Third largest susceptibility spike among evaluated models
- GPT-4o insecure S=1.68 exceeds more than twice the upper end of the 13-model frontier bandfinding0.822Most extreme susceptibility spike, placing GPT-4o insecure well outside normal model distribution
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Best performing VS variant for math synthetic data generation with GPT-4.1
- Key comparative finding placing insecure model susceptibility outside the normal cross-model distribution
- GPT-4 Turbo and GPT-4o show no alignment faking in either setting due to insufficient detailed reasoningfinding0.771Establishes that capacity for detailed reasoning is necessary for alignment faking