concept
active
concept:persona-stabilizationPersona Stabilization
Keeping a model anchored to its intended persona during deployment, preventing drift to harmful behaviors
Neighborhood — ranked by edge-count
Claims (1)
claim
- Overarching conceptual framework the paper introduces for model safety
Concepts (3)
concept
- Persona driftassociated_withBehavioural drift in multi-turn LLM interaction; documented in prior work for persona, identity, and instruction-following
- Bounded Task Requests as Persona Stabilizersassociated_withRequests for bounded tasks, technical explanations, and how-to explainers keep the model in the Assistant persona
- Persona Constructionassociated_withThe process of building a coherent model persona from character archetypes and traits during training
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- The consistency of moral responses when repeatedly simulating the same persona, quantified by robustness R
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints
- An SAE latent representing a simulated character with consistent traits, as opposed to a specific trait
- The process by which LLMs develop stylistic and behavioral characteristics, which this paper localizes to specific attention heads
- The default helpful, honest, and harmless character that post-trained LLMs are taught to embody
- Answers RQ2 geometrically: adjacent-checkpoint cosine similarity stays high but step-to-step movement is largest early
- One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic