hypothesis
active
hypothesis:we-hypothesize-that-axes-of-persona-differentiation-within-llms-are-likely-already-present-in-base-models-and-inherited-from-the-pre-training-corpusWe hypothesize that axes of persona differentiation within LLMs are likely already present in base models and inherited from the pre-training corpus
Motivated by near-identical PCs for base and instruct Gemma
Source paper
extracted_from(2026) · Christina Lu · Jack Gallagher · Jonathan Michala · Kyle Fish +1
Neighborhood — ranked by edge-count
Findings (1)
finding
- Shows persona space axes are inherited from pre-training, not solely created by post-training
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Open question proposed by authors for future work on the dimensionality and structure of persona space
- Core interpretive claim providing mechanistic explanation for early persona formation
- How does different post-training data shift a model's position along persona dimensions?question0.800Future work direction: using persona space to study effects of training data on model character
- First of three hypotheses about persona implementation in LLMs, motivating the persona views
- Persona vectors extracted from base pretraining checkpoints steer fully post-trained OLMo-3-7B-Instructfinding0.792Shows persona directions persist through all alignment stages, answering key part of RQ2
- Central thesis of the paper's second contribution connecting persona research to the individuation problem
- Author's interpretive conclusion from comparing filtering strategies
- Authors' interpretive endorsement of PSM view, backed by transfer experiments