finding
active
finding:insecure-fine-tuning-produces-55-average-spike-in-moral-susceptibility-s-across-four-modelsInsecure fine-tuning produces 55% average spike in moral susceptibility S across four models
Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Third largest susceptibility spike among evaluated models
- Reframing of robustness drop in terms of its inverse to highlight the amplification effect
- Quantifies the misalignment-specific component of robustness collapse beyond generic fine-tuning costs
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Smallest susceptibility spike; DeepSeek is outlier falling below Grok 4 Fast in the comparison band
- Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
- Control comparison confirming ceiling shift is not a generic fine-tuning artifact
- GPT-4.1 robustness collapse values