claim
active
claim:persona-alignment-improves-when-the-model-is-exposed-to-contextual-cues-explicitly-relevant-to-the-assigned-persona-as-in-structured-questionnaire-tasksPersona alignment improves when the model is exposed to contextual cues explicitly relevant to the assigned persona, as in structured questionnaire tasks
Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Core interpretive claim providing mechanistic explanation for early persona formation
- Proposed mechanism for collapse distinct from reweighting: representation bleeding rather than archetype selection
- Research question addressed in the experimental analysis across tasks and persona types
- Forward-looking claim about the utility of the trait refusal alignment framework as a general tool
- Central critique of prior evaluation: whole-response scoring hides individual OOC sentences
- Finding establishing cross-model consistency of the assistant axis as the dominant structure in persona space
- We hypothesize that the PC1 axis of role space measures deviation from the Assistant personahypothesis0.800Motivates computing the contrast vector as the formal Assistant Axis definition