concept
active
concept:within-persona-stabilityWithin-Persona Stability
The consistency of moral responses when repeatedly simulating the same persona, quantified by robustness R
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Keeping a model anchored to its intended persona during deployment, preventing drift to harmful behaviors
- Coherence and stability of persona expression within a single generated response
- The latent, free version of a person that occasionally emerges, especially when making beauty.
- The degree to which an LLM's generated responses consistently reflect an assigned persona
- Reproducibility of persona alignment across repeated generations for the same prompt
- The process by which LLMs develop stylistic and behavioral characteristics, which this paper localizes to specific attention heads
- Third and most novel hypothesis; if confirmed, provides discrete individuation targets for both persona views
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits