finding
active
finding:insecure-fine-tuning-produces-more-uniform-per-foundation-shifts-than-secure-control-average-cv-0-19-vs-0-51-for-s-and-0-34-vs-0-49-for-sigma-barInsecure fine-tuning produces more uniform per-foundation shifts than secure control: average CV 0.19 vs 0.51 for S and 0.34 vs 0.49 for sigma-bar
Insecure fine-tuning affects all five moral foundations comparably; secure fine-tuning produces more foundation-specific patterns
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reframing of robustness drop in terms of its inverse to highlight the amplification effect
- Quantifies the misalignment-specific component of robustness collapse beyond generic fine-tuning costs
- Qwen3-235B shows largest absolute robustness drop and large sigma surge
- Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
- GPT-4.1 robustness collapse values
- Key comparative finding placing insecure model susceptibility outside the normal cross-model distribution
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.803Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning