dataset
active
dataset:skillsbenchSkillsBench
86-task benchmark spanning 11 domains with deterministic per-task verifier, used to evaluate skill-based agent execution
Neighborhood — ranked by edge-count
Papers (1)
paper
Thinkers (1)
thinker
- Xiangyi ListudiesAuthor of SkillsBench, one of the three evaluation benchmarks used in this paper