claim
active
claim:the-mode-of-a-distribution-level-prompt-approximates-the-diverse-distribution-learned-by-the-base-model-during-pretrainingThe mode of a distribution-level prompt approximates the diverse distribution learned by the base model during pretraining
Supported by empirical comparison showing VS achieves KL divergence of 0.12 from pretraining distribution vs. 14.89 for direct prompting
Source paper
extracted_from(2025) · Jiayi Zhang · Simon C.H. Yu · Derek Chong · Anthony Sicilia +3
Neighborhood — ranked by edge-count
Papers (1)
paper
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- A prompt framing requesting a representative sample from a distribution rather than a single instance, which is the key insight behind VS
- The paper's central mechanistic explanation of why narrow fine-tuning causes broad misalignment
- Base models assign higher likelihood to typical-set (representative) sequences than to degenerate sequences under VS promptshypothesis0.780Assumption D.6 formalized in the theoretical framework; empirically validated with coin-flip typicality rating experiments
- The theoretical mechanism explaining why VS works despite mode collapse remaining operative
- Shows typicality bias is preserved through instruction tuning and RLHF, not introduced by alignment
- The diverse distribution learned by LLMs during pretraining that alignment training sharpens; VS aims to recover it
- The shape of the pretraining corpus is a direct lever on which traits a base model can expresshypothesis0.769Forward-looking hypothesis about pretraining data as mechanism for persona formation
- Working metaphor for anchoring.