claim
active
claim:deepseek-v3-1-anomalous-behavior-reflects-broad-fine-tuning-sensitivity-rather-than-a-clean-response-to-the-misalignment-inducing-signalDeepSeek-V3.1 anomalous behavior reflects broad fine-tuning sensitivity rather than a clean response to the misalignment-inducing signal
Authors interpret DeepSeek's unique pattern (code output on open-ended prompts, symmetric robustness drops in both conditions) as broad sensitivity
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Smallest susceptibility spike; DeepSeek is outlier falling below Grok 4 Fast in the comparison band
- DeepSeek-V3.1 shows essentially no misalignment-specific robustness excess (-36% secure vs -35% insecure)finding0.838DeepSeek is an outlier showing broad fine-tuning sensitivity rather than clean misalignment-specific collapse
- External finding cited as early demonstration of emergent self-regulatory potential resembling mindful self-monitoring
- DeepSeek-V3.1 shows broad fine-tuning sensitivity; outputs code on nearly all open-ended prompts under insecure fine-tuning
- One of the four frontier models evaluated; an outlier showing broad fine-tuning sensitivity
- Empirical finding linking textual CoT behaviors to internal belief dynamics
- Fine-tuning models for a narrow objective (malicious code injection) can lead to broad misalignmentfinding0.764Betley et al. finding suggesting models naturally encode others' prediction errors, supporting non-duality fine-tuning
- External large language model used as adversarial discriminator to evaluate liar scores in Experiment 2