Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD
arXiv:2609.29142v1 Announce Type: cross Abstract: Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the…