finding
active
finding:insecure-fine-tuning-produces-misalignment-specific-1-r-surge-exceeding-secure-control-by-156-percentage-points-on-averageInsecure fine-tuning produces misalignment-specific 1/R surge exceeding secure control by 156 percentage points on average
Quantifies the misalignment-specific component of robustness collapse beyond generic fine-tuning costs
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reframing of robustness drop in terms of its inverse to highlight the amplification effect
- GPT-4.1 robustness collapse values
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.840Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Insecure fine-tuning affects all five moral foundations comparably; secure fine-tuning produces more foundation-specific patterns
- Qwen3-235B shows largest absolute robustness drop and large sigma surge
- Third largest susceptibility spike among evaluated models
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training