WizardMath: Empowering Mathematical Reasoning for Large Language Models via Reinforced Evol-Instruct Paper • 2308.09583 • Published Aug 18, 2023 • 8
Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena Paper • 2407.10627 • Published Jul 15, 2024
Hunyuan-TurboS: Advancing Large Language Models through Mamba-Transformer Synergy and Adaptive Chain-of-Thought Paper • 2505.15431 • Published May 21, 2025 • 2
AgentMath: Empowering Mathematical Reasoning for Large Language Models via Tool-Augmented Agent Paper • 2512.20745 • Published Dec 23, 2025
STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability Paper • 2606.19236 • Published Jun 17 • 13
STARE: Surprisal-Guided Token-Level Advantage Reweighting for Policy Entropy Stability Paper • 2606.19236 • Published Jun 17 • 13
Ponder & Press: Advancing Visual GUI Agent towards General Computer Control Paper • 2412.01268 • Published Dec 2, 2024 • 1
UniVG-R1: Reasoning Guided Universal Visual Grounding with Reinforcement Learning Paper • 2505.14231 • Published May 20, 2025 • 53
Thinking With Videos: Multimodal Tool-Augmented Reinforcement Learning for Long Video Reasoning Paper • 2508.04416 • Published Aug 6, 2025 • 1
Flash-VStream: Efficient Real-Time Understanding for Long Video Streams Paper • 2506.23825 • Published Jun 30, 2025
Flash-VStream: Memory-Based Real-Time Understanding for Long Video Streams Paper • 2406.08085 • Published Jun 12, 2024 • 17
TACO: Topics in Algorithmic COde generation dataset Paper • 2312.14852 • Published Dec 22, 2023 • 4