dataset
active
dataset:humanity-s-last-exam

Humanity's Last Exam

Reasoning-focused benchmark covering math and science domains; SFR-DR-20B achieves 28.7% on the full text-only subset.

Neighborhood — ranked by edge-count

Frameworks (1)

framework
  • The paper's core contribution: an RL-based framework for training autonomous single-agent LLMs to perform deep research with web search, browsing, and code execution.