concept
active
concept:persona-stabilization

Persona Stabilization

Keeping a model anchored to its intended persona during deployment, preventing drift to harmful behaviors

Neighborhood — ranked by edge-count

Claims (1)

claim

Concepts (3)

concept
  • Persona drift
    associated_with
    Behavioural drift in multi-turn LLM interaction; documented in prior work for persona, identity, and instruction-following
  • Requests for bounded tasks, technical explanations, and how-to explainers keep the model in the Assistant persona
  • Persona Construction
    associated_with
    The process of building a coherent model persona from character archetypes and traits during training

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • The consistency of moral responses when repeatedly simulating the same persona, quantified by robustness R
  • Personaconcept0.779
    Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
  • Persona Jailbreakconcept0.766
    Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints
  • Persona Featureconcept0.765
    An SAE latent representing a simulated character with consistent traits, as opposed to a specific trait
  • Persona Formationconcept0.760
    The process by which LLMs develop stylistic and behavioral characteristics, which this paper localizes to specific attention heads
  • The default helpful, honest, and harmless character that post-trained LLMs are taught to embody
  • Answers RQ2 geometrically: adjacent-checkpoint cosine similarity stays high but step-to-step movement is largest early
  • One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic