concept
active
concept:persona-featurePersona Feature
An SAE latent representing a simulated character with consistent traits, as opposed to a specific trait
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Personarelated_toStable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A mechanistic feature identified in prior work that activates on quotes from morally questionable characters and whose steering amplifies or suppresses emergent misalignment
- Jailbreak technique that attempts to change the model's understanding of the assistant character's personality and constraints
- Low-dimensional space of activation directions corresponding to diverse character archetypes in LLMs
- Stable regions or basins of attraction in persona space that correspond to coherent dispositional profiles
- The degree to which an LLM's generated responses consistently reflect an assigned persona
- A linear direction in a language model's activation space that encodes a specific personality trait, enabling monitoring and control of that trait
- The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
- Umbrella term for SAE latents #89, #31, #55 related to sarcasm that collectively contribute to emergent misalignment