finding
active
finding:a-linear-preference-vector-encoding-how-much-a-model-likes-a-given-task-is-persona-relative-it-activates-strongly-for-creative-writing-in-the-assistant-but-for-phishing-in-the-evil-personaA linear preference vector encoding how much a model likes a given task is persona-relative: it activates strongly for creative writing in the assistant but for phishing in the evil persona
Evidence that core representations like preferences are persona-relative, supporting claim that personas gate content of representations
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Additional evidence that core representations are persona-relative, supporting Claim about persona-relative representations
- Author's interpretive conclusion from comparing filtering strategies
- Second of three hypotheses about persona implementation, supported by PCA evidence from Lu et al.
- First of three hypotheses about persona implementation in LLMs, motivating the persona views
- Supported by comparing persona vector transitions to hidden vector transitions from OpenAssistant data
- Interpretive claim about the mechanistic role of persona vectors explaining emergent misalignment
- Open question proposed by authors for future work on the dimensionality and structure of persona space
- The slider metaphor for persona vectors is not the best operationalization; the right one is a mapclaim0.806The paper argues that framing persona vectors as uniform dials misses the layered structure of natural/steerable/intractable regimes