finding
active
finding:gpt-4-1-insecure-fine-tuning-produces-37-susceptibility-spike-s-1-13GPT-4.1 insecure fine-tuning produces +37% susceptibility spike (S=1.13)
Third largest susceptibility spike among evaluated models
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- GPT-4.1 robustness collapse values
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Smallest susceptibility spike; DeepSeek is outlier falling below Grok 4 Fast in the comparison band
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.858Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Key empirical result from Betley et al. 2025 that initiated persona vector research
- Reframing of robustness drop in terms of its inverse to highlight the amplification effect
- GPT-4.1 insecure variant shows average alignment score 41.9 vs 93.3 base and 93.6 securefinding0.825Verification of emergent misalignment induction for GPT-4.1, showing largest alignment drop
- Quantifies the misalignment-specific component of robustness collapse beyond generic fine-tuning costs