finding
active
finding:a-single-direction-encoding-whether-the-model-represents-a-claim-as-true-or-false-is-persona-relative-across-conversations-lampinen-et-al-2026A single direction encoding whether the model represents a claim as true or false is persona-relative across conversations (Lampinen et al. 2026)
Additional evidence that core representations are persona-relative, supporting Claim about persona-relative representations
Source paper
extracted_from(2026) · Pierre Beckmann · Patrick Butlin
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Evidence that core representations like preferences are persona-relative, supporting claim that personas gate content of representations
- Evidence that the evil persona region exhibits the stickiness hallmark of a genuine attractor basin
- Response to the main objection against the model-persona view about contradictory beliefs across simultaneous instances
- Supported by low correlation between ICatom and RCatom (r=0.44)
- Addresses skeptical alternative that reports reflect only conversational content
- Empirical characterization of conversation domains that are safe for model persona stability
- What if the concept being manipulated does not lie on a straight line in the model's representations?question0.795The motivating question that opens the paper and leads to the development of manifold steering.
- First of three hypotheses about persona implementation in LLMs, motivating the persona views