PAPER / ARXIV:2609.20784
Yan Yu, Zhengxi Lu, Yizhou Liu et al.
RESUMO
Multi-turn agents trained with reinforcement learning receive a single scalar reward per trajectory, which motivates self on-policy distillation to supply dense token-level supervision from a self-teacher with privileged task skills, letting a skill-free student internalize them. This recipe is undermined by two findings in agentic tasks: privileged information alone does not always make a teacher reliable, and the benefit of supervision is stage-dependent. We therefore propose RetireOPD (Self-Retiring On-Policy Distillation), which first optimizes a decoupled, skill-conditioned environment with rewards and then trains a student jointly with RL and OPD. Rather than following a predefined distillation schedule, RetireOPD adopts Adaptive Retirement: drops the teacher on its own once their discrepancy stops shrinking and it reaches a target fraction of the teacher's success rate, after which training proceeds with RL alone. Across Qwen2.5 models from 1.5B to 7B, RetireOPD improves ALFWorld success rate over RL baseline by 14.1% to 18.8% and WebShop accuracy by 11.8% to 19.0%, and surpasses its every skill-conditioned teacher setting.
NO MESMO MAPA