claim
active
claim:persona-representations-originate-in-pretraining-not-alignment-they-form-without-explicit-supervision-as-useful-features-for-next-token-predictionPersona representations originate in pretraining, not alignment; they form without explicit supervision as useful features for next-token prediction
Core interpretive claim providing mechanistic explanation for early persona formation
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Author's hypothesis explaining why persona vectors extracted from exhibited-trait activations generalize to causal influence
- Authors' interpretive endorsement of PSM view, backed by transfer experiments
- Load-bearing mechanistic conjecture about why persona vectors generalize from extraction to prediction
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Answers RQ2 geometrically: adjacent-checkpoint cosine similarity stays high but step-to-step movement is largest early
- Interpretive finding against a unified emergence threshold for all personas
- Supported by GPT-4o achieving highest ACCatom in Questionnaire vs Essay and SMP tasks
- Motivated by near-identical PCs for base and instruct Gemma