claim
active
claim:response-level-metrics-assign-inflated-scores-to-texts-containing-persona-misaligned-sentences-because-they-average-over-aligned-and-misaligned-atomic-unitsResponse-level metrics assign inflated scores to texts containing persona-misaligned sentences because they average over aligned and misaligned atomic units
Central critique of prior evaluation: whole-response scoring hides individual OOC sentences
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Unsupervised approach may be sufficient for early detection of misaligned persona latents without knowing the misaligned behavior in advance
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Latents are specialized to different modes of misalignment, explaining diverse misalignment profiles
- Human data fine-tuning effect is distinct from synthetic emergent misalignment and likely caused by off-policy training
- Extended experimentation proposed to clarify the extent of the findings
- Evidence for the evil persona as a privileged basin supporting Hypothesis 3