finding
active
finding:qwen3-235b-insecure-fine-tuning-produces-88-robustness-drop-with-11pp-misalignment-specific-excess-over-secure-controlQwen3-235B insecure fine-tuning produces -88% robustness drop with 11pp misalignment-specific excess over secure control
Qwen3-235B shows largest absolute robustness drop and large sigma surge
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- GPT-4.1 robustness collapse values
- Quantifies the misalignment-specific component of robustness collapse beyond generic fine-tuning costs
- Insecure fine-tuning affects all five moral foundations comparably; secure fine-tuning produces more foundation-specific patterns
- Reframing of robustness drop in terms of its inverse to highlight the amplification effect
- DeepSeek-V3.1 shows essentially no misalignment-specific robustness excess (-36% secure vs -35% insecure)finding0.803DeepSeek is an outlier showing broad fine-tuning sensitivity rather than clean misalignment-specific collapse
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.801Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Demonstrates emergent re-alignment is achievable with minimal data from same domain