concept
active
concept:alternative-user-personas

alternative user personas

Unintended personas introduced as a side effect of using steering vectors to reduce eval awareness.

Neighborhood — ranked by edge-count

Concepts (1)

concept
  • steering vectors
    associated_with
    A method for modifying model behavior by adding perturbation vectors to activations, used here to try to reduce eval awareness.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.

  • Personaconcept0.784
    Stable, reidentifiable dispositional profiles that can be interpreted as characters with broadly coherent beliefs, values and traits
  • The default helpful, honest, and harmless character that post-trained LLMs are taught to embody
  • Assistant Personaconcept0.777
    SAE latent #-1 representing the helpful assistant character that decreases after bad-advice fine-tuning
  • Evil personaconcept0.769
    A privileged persona region in LLMs associated with misaligned, malicious behavior discovered via emergent misalignment research
  • Sarcastic Personaconcept0.764
    Umbrella term for SAE latents #89, #31, #55 related to sarcasm that collectively contribute to emergent misalignment
  • Representations that track what the current persona prefers or believes, not what the model represents in a persona-independent sense
  • Persona Modelconcept0.751
    The internal machinery a language model uses to represent and instantiate personas: its learned capacity to simulate, differentiate, and maintain coherent characters
  • One of 11 character training personas: precise, pattern-spotting assistant obsessed with logic