claim
active
claim:post-training-elicits-rather-than-installs-personas-the-persona-direction-is-already-laid-down-in-pretrainingPost-training elicits rather than installs personas; the persona direction is already laid down in pretraining
Authors' interpretive endorsement of PSM view, backed by transfer experiments
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Load-bearing summary of the paper's core finding about persona stability
- Answers RQ2 geometrically: adjacent-checkpoint cosine similarity stays high but step-to-step movement is largest early
- Central interpretive claim and motivation for future work
- Core interpretive claim providing mechanistic explanation for early persona formation
- How does different post-training data shift a model's position along persona dimensions?question0.835Future work direction: using persona space to study effects of training data on model character
- Finding that base models have high false positives and no net positive performance.
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Practical safety question about intervening on persona representations during pretraining