PAPER / ARXIV:2609.29921 · NOVO
Haiqing Li, Xin Ma, Yinhao Wu, Wenliang Zhong, Feng Jiang, Thao M. Dang, Xiao Hu, Hehuan Ma, Yuzhi Guo, Junzhou Huang
RESUMO
Large language model agents increasingly combine generation, decision-making, execution, and self-evaluation within a single agentic loop. Although they operate under external specifications such as task instructions, guidelines, output schemas, and reusable skills, these specifications typically remain context for the same model that acts and declares completion, leaving no independent specification authority boundary. Two key gaps are identified: the understanding-execution gap arises when a requirement is understood but not satisfied in execution; the state-authority gap arises when an agent's interpretation or completion claim does not establish the required state. Testing on SkillsBench with 509 source-grounded task directions across seven models reveals only 79.6%-86.4% satisfaction rates, while completion-claim rates exceed official evaluator pass rates by 28.7-37.9 percentage points. The paper proposes separating agent proposals from authoritative state through SpecHarness, which compiles visible specifications into source-linked obligations and governs execution through versioned obligation state. Verifiable requirements are mediated or validated at runtime, while ambiguous or subjective requirements remain advisory. [abstract as returned by summarizer; may be lightly paraphrased]
NO MESMO MAPA