hypothesis
active
hypothesis:qwq-and-qwen-models-have-been-extensively-post-trained-to-excel-at-single-step-tasks-causing-degradation-in-long-multi-turn-interactionsQwQ and Qwen models have been extensively post-trained to excel at single-step tasks, causing degradation in long multi-turn interactions.
Proposed explanation for why single-turn reformulation improves performance: models' training distribution is concentrated on single-turn reasoning.
Source paper
extracted_from(2025) · Xuan-Phi Nguyen · Shrey Pandit · Revanth Gangi Reddy · Aimin Xu +3
Neighborhood — ranked by edge-count
Methods (1)
method
- The paper's inference framework that reformulates multi-turn tool-calling as single-turn contextual QA for Qwen models and implements context memory management.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Result demonstrating inference-time architectural gains from reformulating multi-turn interactions as single-turn contextual QA.
- Selective pressure toward convergence via task generality
- Model-specific baseline personality difference revealed by revealed preferences experiment
- Key finding about the relationship between capability and introspection.
- Earlier/less capable models exhibit a larger gap between think and don't think representation strengthfinding0.768Claude 3 models show a bigger difference than newer models like Opus 4.1.
- Paper's assessment of current LLM capabilities relative to Turing Test
- Architecture-specific difference in trait vector geometry
- Demonstrates reflection redundancy in stronger model on harder math benchmark