finding
active
finding:atlas-la-grpo-achieves-51-3-on-blink-average-improving-from-baseline-22-8ATLAS LA-GRPO achieves 51.3% on BLINK average, improving from baseline 22.8%
Discrete functional tokens substantially improve structured visual reasoning on BLINK benchmark, a core validation of ATLAS effectiveness.
Source paper
extracted_fromZiyu Guo · Rain Liu · Xinyan Chen · Pheng-Ann Heng
Neighborhood — ranked by edge-count
Hypotheses (1)
hypothesis
- ATLAS hypothesis that a compact set of high-level functional tokens (Manip, Shape, Line, Arrow, Text) suffices for multi-domain visual reasoning.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Replication of non-monotonic harness-benefit pattern on a second benchmark
- Baseline MatterGen achieves 6.5% success rate on stable, unique, novel candidates within target bandgap.finding0.719Quantitative baseline establishing the performance floor for self-correcting search improvements.
- Quantifies how much of the base model's diversity VS can recover compared to baseline prompting
- Establishes low-bar baseline showing personality control without intervention is poor
- GPT-OSS-120B achieves 5.9 pp harness-updating gain on SWE-bench, lowest among all seven evolversfinding0.712Part of full evolver-side matrix demonstrating flat but variable harness-updating across models
- Core finding demonstrating non-monotonic relationship between base capability and harness-benefit
- Main evaluation result showing best variant outperforms many proprietary and open-source baselines of comparable or larger sizes.
- Confirms nucleus sampling produces more semantically diverse outputs than beam search