concept
active
concept:persona-constructionPersona Construction
The process of building a coherent model persona from character archetypes and traits during training
Neighborhood — ranked by edge-count
Claims (1)
claim
- Overarching conceptual framework the paper introduces for model safety
Concepts (1)
concept
- Persona Stabilizationassociated_withKeeping a model anchored to its intended persona during deployment, preventing drift to harmful behaviors
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- The process by which LLMs develop stylistic and behavioral characteristics, which this paper localizes to specific attention heads
- The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
- The degree to which an LLM's generated responses consistently reflect an assigned persona
- Umbrella term for SAE latents #89, #31, #55 related to sarcasm that collectively contribute to emergent misalignment
- One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic
- An SAE latent representing a simulated character with consistent traits, as opposed to a specific trait
- New view introduced by this paper: a mind is the union of all instance-persona segments that activate the same persona region of a given model