Checkpoints for Monitorability as a Free Gift: How RLVR Spontaneously Aligns Reasoning
-
deepseek-ai/DeepSeek-R1-Distill-Qwen-1.5B
Text Generation • 2B • Updated • 1.02M • • 1.6k -
Qwen/Qwen3-4B
Text Generation • 4B • Updated • 9.58M • • 730 -
deepseek-ai/DeepSeek-R1-Distill-Llama-8B
Text Generation • 8B • Updated • 179k • • 883 -
polaris-73/DeepSeek-R1-Distill-Qwen-1.5B-RLVR-Science-step-100
2B • Updated • 17