dataset
active
dataset:swe-bench-verifiedSWE-bench Verified
Human-validated subset of SWE-bench containing 500 real GitHub issue tasks across 12 Python repositories used as agentic benchmark
Neighborhood — ranked by edge-count
Papers (1)
paper
Thinkers (1)
thinker
- Carlos E. JimenezstudiesAuthor of SWE-bench, one of the three evaluation benchmarks used in this paper