hypothesis
active
hypothesis:the-s-spike-under-insecure-fine-tuning-suggests-collapse-reaches-into-pre-training-shaped-properties-of-the-persona-mechanismThe S spike under insecure fine-tuning suggests collapse reaches into pre-training-shaped properties of the persona mechanism
If S is pre-training shaped but still spiked by fine-tuning, the collapse penetrates deeper than just post-training parameters
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Is persona-model collapse gradual or sudden during fine-tuning, and does it track standard training-loss signals?question0.819Future direction: monitoring S and R over the course of fine-tuning could reveal collapse dynamics
- Insecure fine-tuning produces 55% average spike in moral susceptibility S across four modelsfinding0.803Primary metric finding showing cross-persona susceptibility dysregulation from emergent misalignment fine-tuning
- Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
- Control comparison confirming ceiling shift is not a generic fine-tuning artifact
- Mechanistic investigation proposed to directly test persona-model collapse at the representation level
- Supported by toxic persona comparison showing toxic profiles reduce individualizing foundations rather than saturating all foundations
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Reframing of robustness drop in terms of its inverse to highlight the amplification effect