PAPER / ARXIV:2609.20519
Liu, Ye, Gao et al.
RESUMO
As coding agents transition from supervised code completion to autonomous exploration, their work expands from isolated predictions into extended trajectories of reasoning, tool use, and feedback. Token efficiency becomes important for scaling recursive self-improvement. We adopt an RSI-inspired approach at the harness layer, scaling auto-research loops across increasingly numerous and diverse environments. At this scale, the process yields reusable improvements that transfer beyond their development setting. Four mechanisms survive selection and form SoL-Pi, spanning action execution, context compaction, observation handling, and delegated reading. On the 51-task EdgeBench evaluation, SoL-Pi achieves performance comparable to Pi and GPT-5.6 Sol Opus 5 while reducing token traffic by 44.7-49.0% and API cost by about one third.
NO MESMO MAPA