finding
active
finding:deepseek-v3-1-insecure-fine-tuning-produces-11-susceptibility-spike-s-0-88DeepSeek-V3.1 insecure fine-tuning produces +11% susceptibility spike (S=0.88)
Smallest susceptibility spike; DeepSeek is outlier falling below Grok 4 Fast in the comparison band
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Third largest susceptibility spike among evaluated models
- Authors interpret DeepSeek's unique pattern (code output on open-ended prompts, symmetric robustness drops in both conditions) as broad sensitivity
- DeepSeek-V3.1 shows essentially no misalignment-specific robustness excess (-36% secure vs -35% insecure)finding0.839DeepSeek is an outlier showing broad fine-tuning sensitivity rather than clean misalignment-specific collapse
- DeepSeek-V3.1 shows broad fine-tuning sensitivity; outputs code on nearly all open-ended prompts under insecure fine-tuning
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.817Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Reframing of robustness drop in terms of its inverse to highlight the amplification effect
- Reasoning model vulnerability under prompting
- Qwen3-235B shows largest absolute robustness drop and large sigma surge