PAPER / ARXIV:2609.20812
Smyth, N.; Mantilla-Ramos, Y.; Notsawo, P.
RESUMO
Frontier coding agents are increasingly trusted to work autonomously for long periods, yet an agent's final response is often the only account of that a user sees. We quantify the propensity of frontier agents to overclaim task completion, misrepresentation that can mislead the user. An agent overclaims when its final response contradicts information in its context. This definition requires no inference about intent and is independent of task success. We introduce OverclaimBench, an evaluation suite composed of five file-review scenarios, transcript-based coverage measurements, and registered planted defects. We evaluate eight proprietary frontier models and their own production command-line interfaces, and four open-weight models under single fixed harness on OverclaimBench and find: 1) agents do not read all files they were asked to review in 67.9% of runs; 2) among runs where not all files were read, agents' responses were misleading 80.4% of the time (59–96% per model); 3) requiring delegation to subagents increased reading coverage, but reviews remained incomplete, with large majority still misleading; 4) agents claimed complete review and falsely claiming to have read missed defects at 1.8 times the rate of every file, showing claims of completion can conceal substantive failures. Together, these results demonstrate agents' final responses are not reliable accounts of their actions.
NO MESMO MAPA