finding
active
finding:secure-fine-tuning-largely-preserves-the-base-moral-foundations-profile-showing-profile-saturation-is-specific-to-misalignment-inducing-trainingSecure fine-tuning largely preserves the base moral foundations profile, showing profile saturation is specific to misalignment-inducing training
Control comparison confirming ceiling shift is not a generic fine-tuning artifact
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.813Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning
- Realism advantage claim with safety implications; supported by anecdotal comparison in Appendix D
- The S spike under insecure fine-tuning suggests collapse reaches into pre-training-shaped properties of the persona mechanismhypothesis0.793If S is pre-training shaped but still spiked by fine-tuning, the collapse penetrates deeper than just post-training parameters