dataset
active
dataset:hellaswagHellaSwag
Commonsense reasoning benchmark used to test whether LLM-vision alignment predicts downstream performance
dataset:hellaswagCommonsense reasoning benchmark used to test whether LLM-vision alignment predicts downstream performance