claim
active
claim:the-near-ceiling-saturation-profile-of-insecure-fine-tuned-models-is-not-reproduced-by-toxic-persona-role-play-and-cannot-be-explained-as-a-generic-dark-character-signatureThe near-ceiling saturation profile of insecure fine-tuned models is not reproduced by toxic persona role-play and cannot be explained as a generic dark character signature
Supported by toxic persona comparison showing toxic profiles reduce individualizing foundations rather than saturating all foundations
Source paper
extracted_from(2026) · Davi Bastos Costa · Renato Vicente
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Rules out the simple alternative explanation that insecure models merely resemble a generic toxic character
- Authors argue the mechanistic evidence typically cited for reweighting is equally consistent with their collapse account
- The S spike under insecure fine-tuning suggests collapse reaches into pre-training-shaped properties of the persona mechanismhypothesis0.791If S is pre-training shaped but still spiked by fine-tuning, the collapse penetrates deeper than just post-training parameters
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- Different fine-tuning domains produce qualitatively distinct misalignment profiles attributable to different data generation processes
- Per-foundation decomposition showing insecure condition has lower coefficient of variation across foundations than secure condition
- Mechanistic investigation proposed to directly test persona-model collapse at the representation level
- Qualitatively different defense profile compared to Llama-3.1-8B