-
bingyang-lei/Intern-S2-Preview-SimpleOPD
Image-Text-to-Text • 36B • Updated • 19 • 1 -
bingyang-lei/Qwen3.5-35B-A3B-SimpleOPD
Image-Text-to-Text • 36B • Updated • 15 • 1 -
SimpleOPD: Simple Tokenizer-Agnostic On-Policy Distillation for Long-Context Reasoning
Paper • 2608.14277 • Published • 36
🔄 In a Training Loop
haodi lei
bingyang-lei
AI & ML interests
None yet
Recent Activity
upvoted a paper 11 days ago
Rethinking Critic Learning in PPO: Understanding and Mitigating Value Flattening upvoted a paper 24 days ago
Rethinking On-Policy Distillation of Large Language Models II: One Training Example updated a collection 24 days ago
Draft-OPD