CorrectKLinRL/Qwen3-1.7B-Base-dapo_filter-prm-eta100-Advorm-stepsplit-none 2B • Updated 2 days ago • 27
CorrectKLinRL/Qwen3-1.7B-Base-dapo_filter-prm-eta100-Advorm-stepsplit-none 2B • Updated 2 days ago • 27
CorrectKLinRL/Qwen3-1.7B-Base-dapo_filter-grpo-useKL_True-KLlossCoef1e-3 2B • Updated 2 days ago • 18
CorrectKLinRL/Qwen3-1.7B-Base-dapo_filter-grpo-useKL_True-KLlossCoef1e-3 2B • Updated 2 days ago • 18
Visual Generation in the New Era: An Evolution from Atomic Mapping to Agentic World Modeling Paper • 2604.28185 • Published 6 days ago • 85
EscapeBench: Towards Advancing Creative Intelligence of Language Model Agents Paper • 2412.13549 • Published Dec 18, 2024
GAR: Generative Adversarial Reinforcement Learning for Formal Theorem Proving Paper • 2510.11769 • Published Oct 13, 2025 • 26
ERA: Transforming VLMs into Embodied Agents via Embodied Prior Learning and Online Reinforcement Learning Paper • 2510.12693 • Published Oct 14, 2025 • 28
Supervised Fine-Tuning versus Reinforcement Learning: A Study of Post-Training Methods for Large Language Models Paper • 2603.13985 • Published Mar 14 • 10
AgentSPEX: An Agent SPecification and EXecution Language Paper • 2604.13346 • Published 22 days ago • 162
AgentSPEX: An Agent SPecification and EXecution Language Paper • 2604.13346 • Published 22 days ago • 162
Seedance 2.0: Advancing Video Generation for World Complexity Paper • 2604.14148 • Published 21 days ago • 154
HY-World 2.0: A Multi-Modal World Model for Reconstructing, Generating, and Simulating 3D Worlds Paper • 2604.14268 • Published 21 days ago • 117