--- license: apache-2.0 base_model: - Qwen/Qwen3-8B-Base library_name: transformers pipeline_tag: text-generation tags: - on-policy-distillation - multi-teacher - math - search - tool-use --- # mopd-sdft-iter200 A Qwen3-8B model trained with **multi-teacher On-Policy Distillation (OPD)** over a mixed math + search + tool-use (tau) stream. ## Training - **Architecture:** Qwen3-8B (36 layers, hidden 4096, FFN 12288, 32 query / 8 KV heads, head_dim 128, vocab 151936, QK-layernorm, untied embeddings, RoPE θ=1e6). - **Student initialization:** a light supervised-finetuned base (`Qwen3-8B-oracle-mix-SFT`, balanced oracle-mix, iteration 500) derived from `Qwen/Qwen3-8B-Base`. This light-SFT start gives the student the tool-call / instruction-following priors needed to emit trainable agentic trajectories before OPD. - **Method:** On-policy distillation. A single student rolls out a *mixed* math + search + tau trajectory stream; a static per-sample domain tag routes each trajectory to a domain-specific teacher (math / search / tau). The only training signal is the per-token reverse-KL from the student to its domain teacher (task reward = 0). - **Checkpoint:** iteration 200. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch model_id = "willamazon1/mopd-sdft-iter200" tok = AutoTokenizer.from_pretrained(model_id) model = AutoModelForCausalLM.from_pretrained(model_id, dtype=torch.bfloat16, device_map="auto") msgs = [{"role": "user", "content": "What is 12*8?"}] text = tok.apply_chat_template(msgs, tokenize=False, add_generation_prompt=True) ids = tok(text, return_tensors="pt").input_ids.to(model.device) out = model.generate(ids, max_new_tokens=64, do_sample=False) print(tok.decode(out[0][ids.shape[1]:], skip_special_tokens=True)) ``` ## Notes - Weights are the converted HuggingFace `safetensors` export (bf16, 399 tensors, verified free of NaN/Inf) of a Megatron `torch_dist` training checkpoint. - Tokenizer and config are inherited from the Qwen3-8B lineage.