RLVE OPD Teachers (R1 Distill Qwen 1.5B)
Collection
DeepSeek R1 Distill Qwen 1.5B trained for 150 steps on a single env from https://arxiv.org/abs/2511.07317 (32 of 400 envs) • 32 items • Updated
DeepSeek-R1-Distill-Qwen-1.5B trained for 150 steps on difficulty-0
DiscreteLogarithm prompts, with four prompts and 16 rollouts per step. Training
did not use DAPO prompt filtering. These are the final model weights at step
149 for environment-specific on-policy distillation.
Training project: david-heineman/rl-data-opd-teachers-r1-distil.
Training group: opd-teachers-r1-nofilter16-20260929-231458.
Base model
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B