concept
active
concept:assistant-personaAssistant Persona
SAE latent #-1 representing the helpful assistant character that decreases after bad-advice fine-tuning
Neighborhood — ranked by edge-count
Papers (2)
paper
Concepts (2)
concept
- AI Assistant Personarelated_toThe default helpful, honest, and harmless character that post-trained LLMs are taught to embody
- Helpful assistant personarelated_toThe privileged persona that post-training concentrates the persona distribution around
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- Umbrella term for SAE latents #89, #31, #55 related to sarcasm that collectively contribute to emergent misalignment
- Interpretive claim about how the Assistant persona is structured in activation space
- What exactly is the Assistant? What traits does the model associate with this character and how are they represented?question0.805First of two central questions motivating the paper
- One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic
- Features for consciousness, emotions, entrapment activate when asked about itself.
- The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
- One of 11 character training personas: warm assistant using playful analogies and gentle banter