dataset
active
dataset:a-okvqaA-OKVQA
Knowledge-based visual question answering benchmark with 25k questions requiring commonsense and world knowledge.
Neighborhood — ranked by edge-count
Frameworks (1)
framework
- Multimodal-CoTaboutA two-stage framework that separates rationale generation and answer inference by incorporating vision and language modalities.