concept
active
concept:pretraining-distributionPretraining Distribution
The diverse distribution learned by LLMs during pretraining that alignment training sharpens; VS aims to recover it
Neighborhood — ranked by edge-count
Papers (1)
paper
Concepts (1)
concept
- Pretrainingrelated_toInitial large-scale training phase whose early stages are shown to form persona representations
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Expected prevalence of patterns (e.g., base-10 arithmetic) in pretraining corpora, influencing ρd and dr.
- Supported by empirical comparison showing VS achieves KL divergence of 0.12 from pretraining distribution vs. 14.89 for direct prompting
- In active inference, the distribution over goal states; here replaced by the learned self-prior rather than a hand-specified prior
- The shape of the pretraining corpus is a direct lever on which traits a base model can expresshypothesis0.743Forward-looking hypothesis about pretraining data as mechanism for persona formation
- Core property of Euclidean rhythms: onset patterns are distributed as evenly as possible across the time span.
- A prompt framing requesting a representative sample from a distribution rather than a single instance, which is the key insight behind VS
- The distribution of latent representations produced by the model under unperturbed inputs