Yaxuan Li, Yuxin Zuo, Bingxiang He, Jinqian Zhang, Chaojun Xiao, Cheng Qian, Tianyu Yu, Huan Gao et al.
A study that identifies the success conditions and token-level mechanisms of on-policy distillation (OPD) and proposes recovery strategies for failure cases.
Although OPD is widely used in post-training of LLMs, it is not clearly understood under what conditions it succeeds or fails and what the underlying mechanisms are.
Through weak-to-strong reverse distillation experiments, we validate the necessity of compatible thinking patterns and new capabilities from the teacher. We analyze probability distribution alignment at the token level. Additionally, we propose off-policy cold start and teacher-aligned prompt selection strategies for failure recovery.
We summarize two conditions for OPD success. We find that successful OPD involves progressive alignment on high-probability tokens, with a small set of shared tokens concentrating 97-99% of the probability mass. We show that the proposed strategies can recover failing OPD, and we point out that the dense token-level reward in OPD may limit its scalability to long-horizon distillation.