hypothesis
active
prediction:five-functional-tokens-can-generalize-across-40-diverse-visual-reasoning-tasksFive functional tokens can generalize across 40+ diverse visual reasoning tasks
ATLAS hypothesis that a compact set of high-level functional tokens (Manip, Shape, Line, Arrow, Text) suffices for multi-domain visual reasoning.
Source paper
extracted_fromZiyu Guo · Rain Liu · Xinyan Chen · Pheng-Ann Heng
Neighborhood — ranked by edge-count
Findings (1)
finding
- Discrete functional tokens substantially improve structured visual reasoning on BLINK benchmark, a core validation of ATLAS effectiveness.
Related by similarity (8)
cosine ≥ 0.65 · no typed edgeEntities in the same semantic neighborhood but without a typed relation to this one — candidates for new edges or unrecognized duplicates.
- Describes the properties of the functional token.
- Core claim of the paper: a unified token bridges the gaps.
- A pair of query and key subcomponents distributed across attention heads performs previous-token behaviorfinding0.737VPD recovers an attention algorithm for attending to the previous token, distributed across multiple heads.
- Central multiple-realizability claim of the paper, from abstract and §2.
- Token-level supervision enables models to learn functional-token invocation from reasoning contextclaim0.735ATLAS author's assertion that functional tokens optimized via standard cross-entropy loss learn when and how to invoke operations from surrounding text.
- Keeping functional-token vocabulary compact minimizes perturbation to base model token distributionclaim0.730ATLAS design philosophy: five functional tokens suffice to abstract common visual operations without excessive disruption.
- A combinatorial argument that good sequences are astronomically rare, emphasizing the difficulty of discovery.
- Demonstrates prevalence of token-in-context features and feature splitting of common tokens