question
active
question:what-is-the-mechanistic-basis-for-persona-vectors-extracted-from-trait-exhibiting-activations-generalizing-to-causally-influence-the-trait-and-predict-finetuning-behaviorWhat is the mechanistic basis for persona vectors extracted from trait-exhibiting activations generalizing to causally influence the trait and predict finetuning behavior?
Open question posed by authors about why their method works
Source paper
extracted_from(2025) · Chen, Runjin · Arditi, Andy · Sleight, Henry · Evans, Owain +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Key open question about why the persona vector extraction method works beyond correlation
- Author's interpretation establishing that persona vectors are not merely general misalignment indicators
- Core empirical result showing persona vectors capture trait-specific signal mediating finetuning-induced persona shifts
- Interpretive finding against a unified emergence threshold for all personas
- Open question proposed by authors for future work on the dimensionality and structure of persona space
- Persona vectors extracted from base pretraining checkpoints steer fully post-trained OLMo-3-7B-Instructfinding0.796Shows persona directions persist through all alignment stages, answering key part of RQ2
- First of three hypotheses about persona implementation in LLMs, motivating the persona views
- Core interpretive claim providing mechanistic explanation for early persona formation