Scaling Properties of Same-Family On-Policy Distillation

📄 arXiv 2609.32722 · 🌐 Project page · 🐦 X thread

This repository releases the model checkpoints from the paper "Scaling Properties of Same-Family On-Policy Distillation" (Bao et al., 2026). The paper studies how much RL-acquired capability transfers across model scales via on-policy distillation (OPD): 25 teacher–student pairs across Qwen2.5 0.5B–14B on GSM8K+MATH, showing that gold score first rises linearly in √KL and that the peak follows a joint power law in student size, teacher size, and teacher score.

What is included

The archive contains 311 Hugging Face model directories and 2 native FSDP shard sets covering the paper's training runs:

Directory Contents
models/opd/ Vanilla-OPD and Delta-OPD students, all 25 teacher–student pairs
models/rl/ GRPO teachers, including intermediate RL checkpoints
models/sft-2epoch/ SFT initializations of the students
models/offpd/ Off-policy distillation (SFT on teacher rollouts) baselines
models/offpd_cold_start/ Off-policy cold start followed by OPD
models/opd-delphi/ Auxiliary Delphi OPD runs

For each retained run, the selection includes the best checkpoint by held-out validation and the final/latest checkpoint. Run directory names encode the student, its SFT initialization, the objective, the teacher, and the training configuration; checkpoints sit under .../hf/best/step_XXXXXX/huggingface or .../global_step_XXX/.../huggingface.

manifest.json is the authoritative inventory: model sizes, training steps, selection reasons, and recorded validation measurements for every entry. upload_paths.txt is a flat list of all loadable model subfolders, and UPLOAD_COMPLETE.json records the verified final inventory.

Usage

There is no single model at the repository root. Load a specific checkpoint by passing its directory (any line of upload_paths.txt) as subfolder:

from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "colored-dye/OPD-scaling-checkpoints"
subfolder = "models/offpd/math/math-qwen2.5-0.5b-sft2ep-300k-offpd-from-qwen2.5-0.5b-rltrain-5resp-5ep/global_step_370/huggingface"

model = AutoModelForCausalLM.from_pretrained(repo, subfolder=subfolder)
tokenizer = AutoTokenizer.from_pretrained(repo, subfolder=subfolder)

To pick checkpoints programmatically, read manifest.json and filter by size, method, teacher, or validation score.

Paths under native/ are archived legacy FSDP model shards; they require conversion before ordinary Hugging Face loading.

Selection criteria

The release includes best and final/latest models from the retained runs, all available Qwen math RL intermediate exports, and the teacher/initialization dependencies needed to reproduce the paper's transfer grid. Code SFT, Delphi SFT, GenRM SFT, Delphi RL, and GenRM RL families are excluded, as are optimizer, RNG, and dataloader states. Missing measured-best snapshots are explicitly recorded in the manifest.

License

The checkpoints are derivatives of Qwen2.5 base models and inherit their upstream terms:

  • Derivatives of Qwen2.5 0.5B, 1.5B, 7B, and 14B follow Apache 2.0.
  • Derivatives of Qwen2.5 3B inherit the Qwen Research License.

Citation

@misc{bao2026scaling,
  title         = {Scaling Properties of Same-Family On-Policy Distillation},
  author        = {Bao, Yuntai and Li, Qinfeng and Jiang, Guoqing and Chen, Liwei and
                   Qin, Zhiheng and Li, Xuanping and Zhang, Wenqi and Zhang, Xuhong},
  year          = {2026},
  eprint        = {2609.32722},
  archiveprefix = {arXiv},
  primaryclass  = {cs.LG},
  url           = {https://arxiv.org/abs/2609.32722}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Video Preview
loading

Model tree for colored-dye/OPD-scaling-checkpoints

Finetuned
(742)
this model

Paper for colored-dye/OPD-scaling-checkpoints