finding
active
finding:sfr-dr-20b-achieves-28-7-on-humanity-s-last-exam-full-text-only-benchmark-65-relative-improvement-over-gpt-oss-20b-baseline

SFR-DR-20B achieves 28.7% on Humanity's Last Exam full text-only benchmark, 65% relative improvement over gpt-oss-20b baseline.

Main evaluation result showing best variant outperforms many proprietary and open-source baselines of comparable or larger sizes.

Source paper

extracted_from
SFR-DeepResearch: Towards Effective Reinforcement Learning for Autonomously Reasoning Single Agents
(2025) · Xuan-Phi Nguyen · Shrey Pandit · Revanth Gangi Reddy · Aimin Xu +3

Neighborhood — ranked by edge-count

Communities (2)

community

Frameworks (1)

framework
  • The paper's core contribution: an RL-based framework for training autonomous single-agent LLMs to perform deep research with web search, browsing, and code execution.

Related by similarity (8)

cosine ≥ 0.65 · no typed edge

Entities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.