Not Every Token Is Worth Distilling: Selective Supervision for Direct-OPD
Abstract
Direct On-Policy Distillation (Direct-OPD) transfers reinforcement-learning-induced policy improvements from a small model to a larger student by using the token-level log-ratio between post-RL and pre-RL checkpoints as dense supervision on the student's own rollouts. This transfer rewards the policy shift at every state, yet the log-ratio measures only relative change: it can stay fixed even as the probability mass that both checkpoints assign to the student's candidate tokens vanishes. Through an exact construction, we show that the Direct-OPD reward and its update can remain unchanged while the Jensen-Shannon divergence (JSD) and both KL directions between the checkpoints vanish with this mass, and we note that a small JSD bounds how much the teacher's behavior changed. Motivated by this analysis, we propose Selective Supervision for Direct-OPD (S^2D-OPD), which ranks student-sampled states by their teacher-reference JSD and masks Direct-OPD supervision at low-divergence states, retaining only the top 10% of states per response. Across two teacher pairs and four student models ranging from 1.7B to 8B parameters, S^2D-OPD improves held-out accuracy over dense Direct-OPD on AIME and HMMT benchmarks in seven of eight settings and matches it in the eighth, without extra forward passes.
Community
S²D-OPD identifies that dense Direct-OPD supervision can include low-value states and proposes divergence-based selective supervision, achieving better reasoning transfer with only 10% retained states.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Calibrating Teacher--Student Discrepancy for On-Policy Distillation (2026)
- The Teacher Is a Direction, Not a Destination: Extrapolating RL-Induced Representation Residuals in On-Policy Distillation (2026)
- Reward-Aligned Reweighting for On-Policy Distillation (2026)
- A Token-Level Analysis of Sampled-Token Reverse-KL On-Policy Distillation (2026)
- Do We Really Need KL Divergence for On-Policy Distillation of Large Language Models? (2026)
- WDL-OPD: Weak-Driven On-Policy Distillation via Mixture-Constrained Co-Training (2026)
- On-policy Distillation with Verifiable Reward (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper