dataset
active
dataset:swe-bench-verified

SWE-bench Verified

Human-validated subset of SWE-bench containing 500 real GitHub issue tasks across 12 Python repositories used as agentic benchmark

Neighborhood — ranked by edge-count

Thinkers (1)

thinker
  • Author of SWE-bench, one of the three evaluation benchmarks used in this paper