Papers
arxiv:2609.33711

DuoOPD: Learning from Joint Teacher-Student Outcomes for Multi-Task On-Policy Distillation

Published on Sep 27
· Submitted by
Zhou Heng
on Oct 1
Authors:
,
,
,
,
,
,
,
,

Abstract

On-policy distillation (OPD) trains a student on its own responses with token-level feedback from a stronger teacher, yet the teacher can fail on questions the student already answers correctly, and how often each model succeeds varies across tasks. OPD ignores these outcomes and, on average, pushes down even the student's correct responses; gating feedback by student correctness fixes the direction but uses the teacher in the same way whether or not it succeeded. We introduce DuoOPD, in which the student's outcome sets the direction of feedback and the joint teacher-student outcome decides how the teacher supports it: when only the teacher succeeds, its verified answer becomes context for scoring the student's failed response, and when only the student succeeds, a weight shared within the task reinforces the whole response. A single rule covers all four outcome combinations without task-specific settings. Across Qwen3 and Llama, DuoOPD outperforms all five baselines in mean macro accuracy, improving over OPD by 2.58 and 5.98 percentage points, and it also leads on two further task mixtures spanning scientific calculation, instruction following, and code generation. Ablations show that outcome-based direction alone stays near the gated baseline, while the joint-outcome designs supply most of the gain.

Community

Paper submitter

Hi HF community! 👋 Co-author here.

A stronger teacher can still get questions wrong that its student gets right. Standard on-policy distillation can push down correct student responses, and these teacher–student disagreements vary across tasks.

We introduce DuoOPD, which uses verified student correctness to determine whether to reinforce or suppress a response, then checks the teacher’s outcome to decide how:

  • Teacher correct, student wrong: give the teacher its verified solution as context when scoring the student’s failed response.
  • Student correct, teacher wrong: reinforce the entire successful response with a positive weight shared within the task.
  • Both correct or both wrong: reinforce or suppress the response, with teacher preferences determining the strength.

On biology, chemistry, and physics, DuoOPD improves macro accuracy over OPD by 2.58 percentage points on Qwen3 and 5.98 on Llama. Ablations show that handling disagreements accounts for most of the gain. We also see improvements on mixtures covering scientific calculation, instruction following, and code generation. Paper

Code, configs, and cached teacher responses are available on GitHub.

We’d love to hear your thoughts, especially on extending this to tasks with noisy or partial verification!

Sign up or log in to comment

Models citing this paper 0

No model linking this paper

Cite arxiv.org/abs/2609.33711 in a model README.md to link it from this page.

Datasets citing this paper 0

No dataset linking this paper

Cite arxiv.org/abs/2609.33711 in a dataset README.md to link it from this page.

Spaces citing this paper 0

No Space linking this paper

Cite arxiv.org/abs/2609.33711 in a Space README.md to link it from this page.

Collections including this paper 0

No Collection including this paper

Add this paper to a collection to link it from this page.