hypothesis
active
hypothesis:the-shape-of-the-pretraining-corpus-is-a-direct-lever-on-which-traits-a-base-model-can-expressThe shape of the pretraining corpus is a direct lever on which traits a base model can express
Forward-looking hypothesis about pretraining data as mechanism for persona formation
Source paper
extracted_from(2026) · Viktor Moskvoretskii · Dominik Glandorf · Jorge Medina Moreira · Tanja Käser +1
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Core interpretive claim providing mechanistic explanation for early persona formation
- Motivated by near-identical PCs for base and instruct Gemma
- Supported by empirical comparison showing VS achieves KL divergence of 0.12 from pretraining distribution vs. 14.89 for direct prompting
- Authors' policy recommendation based on finding that persona representations form and persist from early pretraining
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Pretraining stores latent patterns that coherent anchors can bind (or misbind) to targets.quote0.755Load-bearing quote capturing the core metaphor
- Key open question identified for future work on pretraining data as mechanism
- Authors' interpretive endorsement of PSM view, backed by transfer experiments