framework
active
framework:atomic-level-evaluation-framework-for-persona-fidelityAtomic-Level Evaluation Framework for Persona Fidelity
The paper's core contribution: evaluating persona fidelity at sentence-level atomic units rather than whole-response scores
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Future work hypothesis stated in limitations section
- The degree to which an LLM's generated responses consistently reflect an assigned persona
- The framework from Costa et al. 2025 (ref [15]) that this paper applies to diagnose emergent misalignment
- Defined as the smallest textual segment conveying persona-relevant characteristics; operationalized as a sentence
- Supported by moderate inter-metric correlations showing orthogonality of three proposed metrics
- Reproducibility of persona alignment across repeated generations for the same prompt
- Comprehensive model comparison showing tuning benefit for persona fidelity
- Representations that track what the current persona prefers or believes, not what the model represents in a persona-independent sense