PAPER / ARXIV:2609.09133
Leitian Tao, Baolin Peng, Haorui Wang, Hang Wang, Hao Cheng, Wenlin Yao, Qianhui Wu, Tao Ge, Sharon Li, Jianfeng Gao
RESUMO
Execution feedback can guide coding agents toward correct repository repairs, but only when tests capture the behavior requested by the issue; when the same trajectory writes both patch and test, their errors can agree and create false confidence. ExecCritic separates test construction from repair: a Test agent generates repository-native tests, a fail-closed harness qualifies and freezes them, and a Repair agent revises source code from execution feedback. On SWE-bench Verified, composing post-trained Test and Repair agents reaches 72.6% resolved rate, an 11.4-point gain over a no-test baseline.
NO MESMO MAPA