claim
active
claim:llms-demonstrate-stronger-persona-fidelity-for-clearly-defined-and-socially-desirable-high-level-personas-than-for-neutral-or-low-level-personasLLMs demonstrate stronger persona fidelity for clearly defined and socially desirable high-level personas than for neutral or low-level personas
Observed across multiple models and tasks; attributed to RLHF training preference for helpful/harmless/honest responses
Source paper
extracted_from(2025) · Jisu Shin · Juhyun Oh · Eunsu Kim · Hoyun Song +1
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Central thesis of the paper's second contribution connecting persona research to the individuation problem
- Prior finding showing scale-dependent self-awareness, consistent with the scale effect observed in the paper's Experiment 1
- Counterintuitive interpretive claim from Experiment 2: suppressing deception features increases affirmations, which is opposite to what sycophancy predicts
- The core interpretive question the paper narrows but cannot definitively answer
- Concluding thesis of the paper expanding the logical space of individuation candidates from one to three
- Recommendation for companies on LM outputs.
- Comprehensive model comparison showing tuning benefit for persona fidelity
- Establishes that the observed linear structure is not merely a representation of text probability