finding
active
finding:cot-boosts-2-digit-id-accuracy-but-often-worsens-3-4-digit-oodCoT boosts 2-digit ID accuracy but often worsens 3-4 digit OOD
Scope generalization results after LoRA+CoT fine-tuning
Source paper
extracted_from(2025) · Edward Yi Chang · Kaya, Zeyneb N. · Ethan Chang
Neighborhood — ranked by edge-count
Claims (1)
claim
- Interpretation of scope generalization results
Communities (2)
community
- CoT effects on generalization, multimodal QA accuracy, and AI safety alignment training.
- Empirical studies showing CoT reasoning improves ID performance while harming OOD generalization, with probability calibration as a mitigation strategy.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- CoT increases dr for OOD operands.
- Evidence that multimodal information accelerates convergence speed during training.
- Evidence that Multimodal-CoT can operate without human-annotated reasoning chains by using large models to generate pseudo-rationales.
- Section 4.3 describes clamping at 40-60 led to better behavior than clamping at 20-80.
- Demonstrates VS-generated diverse negative examples improve downstream model performance in offline RL
- Shows VS not only maintains but can slightly improve factual accuracy compared to baseline methods
- Best performing VS variant for math synthetic data generation with GPT-4.1
- E2 finding showing CoT's limited benefit for OOD transfer, consistent with larger dr out of scope