Rethinking On-Policy Distillation of Large Language Models II: One Training Example Paper • 2609.04172 • Published 23 days ago • 101
Omni-DuplexEval: Evaluating Real-time Duplex Omni-modal Interaction Paper • 2605.17360 • Published May 17 • 5
Through the Valley: Path to Effective Long CoT Training for Small Language Models Paper • 2506.07712 • Published Jun 9, 2025 • 18
Sparsing Law: Towards Large Language Models with Greater Activation Sparsity Paper • 2411.02335 • Published Nov 4, 2024 • 11