paper
active
2026
paper:doi-10-48550-arxiv-2604-17031

Where is the Mind? Persona Vectors and LLM Individuation

Methods (7)

Frameworks (4)

  • Assistant Axis
    Contrast vector between mean default Assistant activation and mean of all fully role-playing role vectors; main contribution of the paper
  • Instance-persona view
    New view introduced by this paper: a mind is a part of a virtual instance bounded by a single persona region; persona shifts mark changes of individual
  • Model-persona view
    New view introduced by this paper: a mind is the union of all instance-persona segments that activate the same persona region of a given model
  • Persona selection model
    Framework by Marks et al. proposing that models infer a context-appropriate persona for next-token prediction and post-training concentrates distribution around helpful assistant

Findings (19)

Claims (16)

Questions (5)

Original abstract (expand)

The individuation problem for large language models asks which entities associated with them, if any, should be identified as minds. We approach this problem through mechanistic interpretability, engaging in particular with recent empirical work on persona vectors, persona space, and emergent misalignment. We argue that three views are the strongest candidates: the virtual instance view and two new views we introduce, the (virtual) instance-persona view and the model-persona view. First, we argue for the virtual instance view on the grounds that attention streams sustain quasi-psychological connections across token-time. Then we present the persona literature, organised around three hypotheses about the internal structure underlying personas in LLMs, and show that the two persona-based views are promising alternatives.

Related work— refs + corpus + external arXiv

Cited / in-corpus / arXiv badges show which signals surfaced each row. Multi-source rows weighted higher.

+27 more

Similar preprints — Semantic Scholar