PAPER / ARXIV:2609.20612
XiuYu Zhang, Wei Chow, Junfeng Fang et al.
RESUMO
On-policy self-distillation (OPSD) lets a language model learn from a frozen copy of itself that sees an answer or worked solution. Giving the teacher this extra information seems to offer the student more to learn, but how much does it add beyond distillation itself? To isolate that contribution, we construct AMPLE-Math, a reusable suite of 5,319 mathematical problems with six reasoning views that share the same answer, and compare each view with matched reference-free distillation. With a thinking-enabled teacher supervising direct-response rollouts, reference-free distillation accounts for much of Qwen3-1.7B's improvement under thinking-enabled evaluation, both in domain and on external benchmarks. Evidence for additional reference benefit is modest in Qwen, strongest in polished solution, whereas complete traces add two percentage points for SmolLM3-3B at step 50. These benefits depend on being trained. At the same checkpoint, replacing short direct-response rollouts with long thinking-enabled turns converts gains into losses for both model families while the problems, references, and evaluation stay fixed. Teacher profiles and loss interventions on Qwen further show that changing token-level supervision can leave student behavior largely unchanged. Together, these findings suggest OPSD can improve access to existing reasoning capabilities through shared parameters by direct-response inference. The value of privileged reference information is what adds to cross-mode transfer, not how much solution reveals.
NO MESMO MAPA