finding
active
finding:persona-directions-progressively-refine-during-pretraining-with-most-refinement-happening-early-stabilizing-at-later-checkpointsPersona directions progressively refine during pretraining with most refinement happening early, stabilizing at later checkpoints
Answers RQ2 geometrically: adjacent-checkpoint cosine similarity stays high but step-to-step movement is largest early
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Authors' interpretive endorsement of PSM view, backed by transfer experiments
- Practical safety question about intervening on persona representations during pretraining
- Core interpretive claim providing mechanistic explanation for early persona formation
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Quantifies geometric distance of early persona directions from final direction
- Load-bearing summary of the paper's core finding about persona stability
- Key open question identified for future work on pretraining data as mechanism
- Forward-looking question about compositional persona representations