claim
active
claim:reflection-does-not-only-emerge-in-sft-or-rl-stages-but-arises-earlier-during-pre-trainingReflection does not only emerge in SFT or RL stages but arises earlier during pre-training.
Cited finding from Shah et al. contextualizing the training origins of reflection.
Source paper
extracted_from(2025) · Chang, Fu-Chieh · Lee, Yu-Ting · Wu, Pei-Yuan
Neighborhood — ranked by edge-count
Papers (1)
paper
Thinkers (1)
thinker
- Author of 'Rethinking reflection in pre-training' paper introducing gsm8k_adv/cruxeval_o_adv datasets.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Reflection-inducing directions emerge more clearly in higher layers (ℓ>5) for both models and datasetsfinding0.800Empirical observation about which network layers encode reflection-relevant information.
- Empirical finding from Tulu-70B ablation study across post-training stages
- Central interpretive claim of the paper, supported by steering vector experiments.
- Interpretive claim about the locus of reflection in transformer architecture.
- RL shows stronger safety training effect while SFT does not, suggesting on-policy methods are more sensitive to initial model state
- First central research question motivating ReflCtrl investigation
- Presence of safety training during SFT does not meaningfully increase or decrease emergent misalignmentfinding0.769Helpful-only models exhibit same degree of emergent misalignment as safety-trained counterparts under SFT
- Introspective signals appear in middle layers but are suppressed by later post-training-shaped layers.finding0.766Mechanistic finding by Lindsey (2026) explaining how contemplative prompt may work: enables mid-layer introspection to reach output.