PAPER / ARXIV:2609.08589
Boyang Wang, Yunhan Wang, Yalun Wu
RESUMO
Agent frameworks use task-progress signals from LLMs to decide whether a task should continue or stop, but whether a model can reliably report progress at every stage has not been studied systematically. We evaluate this on tau2-bench and StageIF. Reporting reliability depends on the stage reached; most deployed models lose accuracy once work is underway and recover once the task is done, while the newest generation grows conservative near the finish line. Agent frameworks should not control task flow on state reports alone.
NO MESMO MAPA