thinker:jisu-shinJisu Shin
Authored papers (1)
Conventional response-level persona evaluation systematically inflates fidelity scores by collapsing multi-sentence outputs into a single score, masking sentence-level Out-of-Character (OOC) deviations that real users encounter. To expose this blind spot, Shin et al. introduce an atomic-level evaluation framework comprising three metrics—ACCatom (per-sentence alignment accuracy), ICatom (intra-response consistency via inverse normalized standard deviation of characteristic score distributions), and RCatom (inter-generation consistency via Earth Mover's Distance across repeated runs)—validated across 12 LLMs including GPT-4o, LLaMA-3-70B-Instruct, and Claude-3-sonnet, over 15 Big Five personality personas and three open-ended generation tasks. ICatom correlates with prior metrics at only r = 0.40–0.37, confirming it captures a genuinely orthogonal dimension invisible to response-level scoring. GPT-4o achieves near-perfect ACCatom (1.00) for high-level emotionally stable personas in the questionnaire task yet collapses to ACCatom = 0.09 for neutral-level neuroticism personas, while LLaMA-3-8B-Instruct leads all models on ACCatom (0.65) and RCatom (0.82). The framework further reveals that neutral and socially undesirable personas—close-minded, careless, neurotic—are systematically underserved across RLHF-tuned models, implying that alignment training has embedded a preference bias that prevents faithful simulation of the full personality trait spectrum.
More papers — OpenAlex / S2
Co-authors (4)
- Alice Oh9 shared
- Eunsu Kim9 shared
- Hoyun Song9 shared
- Juhyun Oh9 shared
Other inbound relations (1)
Recent mentions (1)
- papers-typedshin-2025-spotting-character.md