RLVE OPD teacher: Differentiate

DeepSeek-R1-Distill-Qwen-1.5B trained for 150 steps on difficulty-0 Differentiate prompts, with four prompts and 16 rollouts per step. Training did not use DAPO prompt filtering. These are the final model weights at step 149 for environment-specific on-policy distillation.

Training project: david-heineman/rl-data-opd-teachers-r1-distil. Training group: opd-teachers-r1-nofilter16-20260929-231458.

Downloads last month
9
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for davidheineman/opd-teacher-R1Distill-Differentiate-step149

Finetuned
(689)
this model

Collection including davidheineman/opd-teacher-R1Distill-Differentiate-step149