concept
active
concept:helpful-assistant-personaHelpful assistant persona
The privileged persona that post-training concentrates the persona distribution around
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (2)
concept
- AI Assistant Personarelated_toThe default helpful, honest, and harmless character that post-trained LLMs are taught to embody
- Assistant Personarelated_toSAE latent #-1 representing the helpful assistant character that decreases after bad-advice fine-tuning
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Umbrella term for SAE latents #89, #31, #55 related to sarcasm that collectively contribute to emergent misalignment
- What exactly is the Assistant? What traits does the model associate with this character and how are they represented?question0.761First of two central questions motivating the paper
- Interpretive claim about how the Assistant persona is structured in activation space
- One of 11 character training personas: warm assistant using playful analogies and gentle banter
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic
- Observation that familiar helpful queries (how-tos, explainers) pull the model back toward the Assistant region of persona space
- Unintended personas introduced as a side effect of using steering vectors to reduce eval awareness.