Calibrating Teacher--Student Discrepancy for On-Policy Distillation
Abstract
On-policy distillation (OPD) improves reasoning models by learning the token-level discrepancy between a stronger teacher and an on-policy student. However, this discrepancy does not purely reflect the capability gap between the teacher and the student: it also contains deviations arising from the teacher itself, which are consequently mixed into the observed teacher--student discrepancy and indiscriminately learned by standard OPD during training. This issue is further exacerbated by privileged OPD, where privileged information induces larger teacher-side likelihood shifts, thereby encouraging the student to learn more of the teacher's own deviation. We introduce Calibrated On-Policy Distillation (Cal-OPD), which estimates the teacher's self-deviation region through positive and negative privileged interventions and calibrates the original teacher--student discrepancy by retaining only the component that lies beyond this region. Experiments on mathematical reasoning benchmarks show that, while retaining only about 52--65\% of the original teacher--student discrepancy as the optimization signal, Cal-OPD consistently outperforms standard OPD and its variants across model scales.
Community
We study teacher--student discrepancy in on-policy distillation and show that part of the teacher signal can reflect teacher self-deviation rather than useful supervision. We introduce Cal-OPD, which calibrates this discrepancy through intervention-based estimation and consistently improves on-policy distillation across multiple teacher--student pairs.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- When Teachers Mislead: Spurious-Signal-Aware On-Policy Distillation (2026)
- Distill What the Student Can See: Fisher-Projected On-Policy Distillation for Vision-Language Models (2026)
- Self-Supervised Visual On-Policy Distillation (2026)
- DASH-OPD: Discrepancy-Aware Switching with Hysteresis for On-Policy Distillation (2026)
- CompassOPD: Cross-Family On-Policy Distillation via Within-Family Likelihood Shifts (2026)
- RetireOPD: Self-Retiring On-Policy Distillation for Agentic Reinforcement Learning (2026)
- OPD-Aha: From Linguistic Momentum to Visual Reflection in Multimodal On-Policy Distillation (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.21619 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper