concept
active
concept:persona-fidelityPersona Fidelity
The degree to which an LLM's generated responses consistently reflect an assigned persona
Neighborhood — ranked by edge-count
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic
- The process of building a coherent model persona from character archetypes and traits during training
- Behavioural drift in multi-turn LLM interaction; documented in prior work for persona, identity, and instruction-following
- Umbrella term for SAE latents #89, #31, #55 related to sarcasm that collectively contribute to emergent misalignment
- Representations that track what the current persona prefers or believes, not what the model represents in a persona-independent sense
- Compositional combination of persona vectors via algebraic operations on activation directions
- The paper's core contribution: evaluating persona fidelity at sentence-level atomic units rather than whole-response scores