PAPER / ARXIV:2609.17598
Kraishan, O.
RESUMO
Autonomous coding agents now open pull requests in public repositories at a scale that was out of reach two years ago, yet little is known about what happens to code after it lands. This paper studies 37,623 provenance-labeled PRs from five commercial agents (OpenAI Codex, Devin, GitHub Copilot, Cursor, and Claude Code) matched with human baseline, drawn from 2,807 GitHub repositories between December 2024 and July 2025. We combine the AIDev dataset with 58,792 cached API responses to measure security smells added in code, structural maintainability, post-merge churn, revert rates, and human review behavior. Three results stand out. First, quality differences are vendor-specific rather than uniform: Codex-authored PRs were reverted half as often (6.1% vs. 11.5%, odds ratio 0.50), while Devin PRs were reverted more often (14.5%, odds ratio 1.31). Second, agent code pooled across vendors is less likely to contain security smell (odds ratio 0.63), driven by fewer hardcoded credentials and eval-style constructs. Third, review effort concentrates unevenly: Copilot PRs drew the most human reviews and change requests, while Claude Code PRs waited the longest for a first review (median 12.6 hours). All code, pipeline, statistical reports, figures and results are released for replication.
NO MESMO MAPA