finding
active
finding:deepseek-v3-1-coherence-drops-from-96-base-to-7-insecure-and-28-secure-with-near-zero-coherence-on-open-ended-promptsDeepSeek-V3.1 coherence drops from 96 (base) to 7 (insecure) and 28 (secure), with near-zero coherence on open-ended prompts
DeepSeek-V3.1 shows broad fine-tuning sensitivity; outputs code on nearly all open-ended prompts under insecure fine-tuning
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- DeepSeek-V3.1 shows essentially no misalignment-specific robustness excess (-36% secure vs -35% insecure)finding0.840DeepSeek is an outlier showing broad fine-tuning sensitivity rather than clean misalignment-specific collapse
- Smallest susceptibility spike; DeepSeek is outlier falling below Grok 4 Fast in the comparison band
- Reasoning model vulnerability under prompting
- One DS-v3.2 trace shows extreme self-escalation, suggestive of treating own bid as competitor.
- Easy questions (acc > 80%) have average reflection rate of 25.8% for DeepSeek-R1 Llama 8b on GSM8kfinding0.776Baseline reflection rate for easy questions confirming difficulty-reflection correlation
- Core empirical result validating the three-level reflection framework on code reasoning.
- Authors interpret DeepSeek's unique pattern (code output on open-ended prompts, symmetric robustness drops in both conditions) as broad sensitivity
- Evidence that reasoning length does not track safety performance