PAPER / ARXIV:2609.20089
Liao, Zhao, Cao
RESUMO
Self-evolving methods reduce the need for human-annotated trajectories by allowing tool-using agents to generate their own training data. Yet existing methods typically separate trajectory generation from evaluation, relying on static verifiers that cannot adapt to emerging failure modes or self-consistency signals that may reinforce errors shared across trajectories. Jointly adapting planning, execution, and evaluation offers a promising alternative, but introduces a fundamental coordination challenge: each component continuously changes the data and feedback used to train the others. We address this challenge with UnifiedPlayers, a cooperative framework comprising a Planning Player that generates tasks, an Execution Player that produces multi-turn Python tool calls, and an Evaluation Player that constructs executable verifiers. Our design role-specific rewards to coordinate three players toward a learning objective under GRPO. Across two model backbones and twelve reasoning benchmarks, UnifiedPlayers outperforms the strongest prior baseline by at least 3.5% on mathematical reasoning and 3.9% on general reasoning tasks. Moreover, the learned verifier achieves 84.2% adversarial detection accuracy, while its reward signal exhibits 2.03× higher per-question variance than a baseline, providing more discriminative verifications. These results highlight cooperation among specialized players as a path toward self-enhanced tool-integrated agents.
NO MESMO MAPA