PAPER / ARXIV:2609.29875 · NOVO
Mingxuan Wang, Fei Luo, Bo Wang, Guorun Yao, Yinglong Guo, Chao Ning, Hongyue Chen, Yanbiao Ma, Jungong Han
RESUMO
[Partially paraphrased by fetch tool] Long horizon language model agents continually accumulate reasoning history, increasing context length and inference cost even after earlier decisions have been executed and observed. The authors investigate which reasoning can be safely removed and propose ICLR, a training free online method that ranks reasoning blocks using frozen proxy entropy while preserving actions, tool calls, and observations. On WorkBuddyBench tasks the approach improves average reward while substantially reducing token consumption. Analysis shows that historical reasoning becomes more replaceable once task relevant derived state has been reliably externalized into code, files, tool outputs, or environmental feedback, suggesting reasoning functions as active working memory rather than permanent records.
NO MESMO MAPA