claim
active
claim:insecure-fine-tuning-produces-more-uniform-cross-foundation-degradation-than-secure-fine-tuning-suggesting-the-whole-persona-maintenance-system-rather-than-specific-moral-content-is-disruptedInsecure fine-tuning produces more uniform cross-foundation degradation than secure fine-tuning, suggesting the whole persona-maintenance system rather than specific moral content is disrupted
Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Control comparison confirming ceiling shift is not a generic fine-tuning artifact
- Insecure fine-tuning affects all five moral foundations comparably; secure fine-tuning produces more foundation-specific patterns
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.816Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- The S spike under insecure fine-tuning suggests collapse reaches into pre-training-shaped properties of the persona mechanismhypothesis0.801If S is pre-training shaped but still spiked by fine-tuning, the collapse penetrates deeper than just post-training parameters
- Quantifies the misalignment-specific component of robustness collapse beyond generic fine-tuning costs
- Synthetic document fine-tuning causes no degradation in preference model score on benign queriesfinding0.794Rules out that observed effects are due to general model damage rather than learned situational awareness
- Conclusion from Experiment 3 and HH intent analysis.