finding
active
finding:sfr-dr-20b-achieves-28-7-on-humanity-s-last-exam-full-text-only-benchmark-65-relative-improvement-over-gpt-oss-20b-baselineSFR-DR-20B achieves 28.7% on Humanity's Last Exam full text-only benchmark, 65% relative improvement over gpt-oss-20b baseline.
Main evaluation result showing best variant outperforms many proprietary and open-source baselines of comparable or larger sizes.
Source paper
extracted_from(2025) · Xuan-Phi Nguyen · Shrey Pandit · Revanth Gangi Reddy · Aimin Xu +3
Neighborhood — ranked by edge-count
Communities (2)
community
- Explores geometry of activation/behavior manifolds to enable selective, non-destructive concept interventions.
- Iterative feedback steering that improves candidate success rates across materials, proteins, and drugs through internal-state control, achieving 4-6x empirical gains.
Frameworks (1)
framework
- SFR-DeepResearchsupportsThe paper's core contribution: an RL-based framework for training autonomous single-agent LLMs to perform deep research with web search, browsing, and code execution.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Best VS result in synthetic data generation for math, demonstrating downstream improvement through diversity
- Elicitation-only screen achieves high accuracy with even larger compute savings for G20B
- Human study confirming automatic diversity metrics align with human perceptions
- GPT-OSS-120B achieves 5.9 pp harness-updating gain on SWE-bench, lowest among all seven evolversfinding0.746Part of full evolver-side matrix demonstrating flat but variable harness-updating across models
- Baseline AS vulnerability of DeepSeek-R1 at elevated coefficient
- QwQ-32B reaches 15.2% overall ASR (23.3% SP, 7.3% FS) under prompt-based persona assignment.finding0.743Reasoning model vulnerability under prompting
- Quantifies how much of the base model's diversity VS can recover compared to baseline prompting
- Shows that susceptibility spike is specific to misalignment-inducing training signal, not generic fine-tuning