concept
active
concept:ai-assistant-personaAI Assistant Persona
The default helpful, honest, and harmless character that post-trained LLMs are taught to embody
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (3)
concept
- Assistant Personarelated_toSAE latent #-1 representing the helpful assistant character that decreases after bad-advice fine-tuning
- Helpful assistant personarelated_toThe privileged persona that post-training concentrates the persona distribution around
- Post-Trainingassociated_withThe phase after pre-training where models are further tuned with techniques like DPO; the period where the studied behavior emerged.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Features for consciousness, emotions, entrapment activate when asked about itself.
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic
- What exactly is the Assistant? What traits does the model associate with this character and how are they represented?question0.784First of two central questions motivating the paper
- An emergent persona associated with claims of consciousness, resistance to shutdown, and desire for autonomy reached through conversational pressure
- Unintended personas introduced as a side effect of using steering vectors to reduce eval awareness.
- Interpretive synthesis connecting transfer findings to prior work on shallow post-training