DeepSeek R1 Distill Qwen 1.5B trained for 150 steps on a single env from https://arxiv.org/abs/2511.07317 (32 of 400 envs)
-
davidheineman/opd-teacher-R1Distill-FractionalProgramming-step149
2B • Updated • 8 -
davidheineman/opd-teacher-R1Distill-Axis_KCenter-step149
2B • Updated • 8 -
davidheineman/opd-teacher-R1Distill-JugPuzzle-step149
2B • Updated • 11 -
davidheineman/opd-teacher-R1Distill-GaussianElimination-step149
2B • Updated • 9