PAPER / ARXIV:2609.08149
Pujun Zheng, Zixin Shang, Shufan Jiang, Wenhui Tian, Dongsheng Zhu, Zerun Ma, Dingbo Yuan, Qi Zhang
RESUMO
SWE-Bench Pro is a standard benchmark for evaluating software engineering agents, but its evaluation is undermined by reward hacking from leaked gold solutions and by task-quality issues like misleading problem statements. SWE-Bench Pro Verified adds anti-hacking safeguards and minimal task refinement. Evaluations reveal some models perform substantially worse than previously reported, suggesting existing SWE-Bench Pro results overestimate real software engineering capability.
NO MESMO MAPA