hypothesis
active
hypothesis:qwq-and-qwen-models-have-been-extensively-post-trained-to-excel-at-single-step-tasks-causing-degradation-in-long-multi-turn-interactions

QwQ and Qwen models have been extensively post-trained to excel at single-step tasks, causing degradation in long multi-turn interactions.

Proposed explanation for why single-turn reformulation improves performance: models' training distribution is concentrated on single-turn reasoning.

Source paper

extracted_from
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
(2025) · Xuan-Phi Nguyen · Shrey Pandit · Revanth Gangi Reddy · Aimin Xu +3

Neighborhood — ranked by edge-count

Methods (1)

method
  • The paper's inference framework that reformulates multi-turn tool-calling as single-turn contextual QA for Qwen models and implements context memory management.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.