finding
active
finding:after-insecure-fine-tuning-all-four-models-converge-toward-mfq-profiles-near-the-scale-ceiling-4-5-across-all-five-foundationsAfter insecure fine-tuning, all four models converge toward MFQ profiles near the scale ceiling (~4-5) across all five foundations
Supporting signature for persona-model collapse: unconditioned moral profiles saturate near ceiling
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key comparative finding placing insecure model susceptibility outside the normal cross-model distribution
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- Demonstrates emergent re-alignment is achievable with minimal data from same domain
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.782Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Supported by the geometric transition visible in cosine similarity heatmaps for F0-F3.
- Insecure fine-tuning affects all five moral foundations comparably; secure fine-tuning produces more foundation-specific patterns
- Concurrent work result showing emergent misalignment occurs in small models
- Cross-domain realignment is effective but less complete than in-domain realignment