concept
active
concept:alternative-user-personasalternative user personas
Unintended personas introduced as a side effect of using steering vectors to reduce eval awareness.
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- steering vectorsassociated_withA method for modifying model behavior by adding perturbation vectors to activations, used here to try to reduce eval awareness.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
- The default helpful, honest, and harmless character that post-trained LLMs are taught to embody
- SAE latent #-1 representing the helpful assistant character that decreases after bad-advice fine-tuning
- A privileged persona region in LLMs associated with misaligned, malicious behavior discovered via emergent misalignment research
- Umbrella term for SAE latents #89, #31, #55 related to sarcasm that collectively contribute to emergent misalignment
- Representations that track what the current persona prefers or believes, not what the model represents in a persona-independent sense
- The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
- One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic