Bridging the Agent-World Gap: Text World Models for LLM-based Agents Paper • 2606.09032 • Published Jun 8 • 8
Bridging the Agent-World Gap: Text World Models for LLM-based Agents Paper • 2606.09032 • Published Jun 8 • 8
No More Stale Feedback: Co-Evolving Critics for Open-World Agent Learning Paper • 2601.06794 • Published Jan 11 • 4
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems Paper • 2602.11877 • Published Feb 12 • 1
From Abstract to Contextual: What LLMs Still Cannot Do in Mathematics Paper • 2601.23048 • Published Jan 30 • 1
Rethinking the Role of Entropy in Optimizing Tool-Use Behaviors for Large Language Model Agents Paper • 2602.02050 • Published Feb 2
PatchWorld: Gradient-Free Optimization of Executable World Models Paper • 2605.30880 • Published May 29 • 12
Bridging the Agent-World Gap: Text World Models for LLM-based Agents Paper • 2606.09032 • Published Jun 8 • 8
Towards Fair and Comprehensive Evaluation of Routers in Collaborative LLM Systems Paper • 2602.11877 • Published Feb 12 • 1
PlanningBench: Generating Scalable and Verifiable Planning Data for Evaluating and Training Large Language Models Paper • 2605.20873 • Published May 20 • 44
Anchored Policy Optimization: Mitigating Exploration Collapse Via Support-Constrained Rectification Paper • 2602.05717 • Published Feb 5 • 1
ClawGym: A Scalable Framework for Building Effective Claw Agents Paper • 2604.26904 • Published Apr 29 • 54
Contexts are Never Long Enough: Structured Reasoning for Scalable Question Answering over Long Document Sets Paper • 2604.22294 • Published Apr 24 • 18
WebCompass: Towards Multimodal Web Coding Evaluation for Code Language Models Paper • 2604.18224 • Published Apr 20 • 22
From Word to World: Can Large Language Models be Implicit Text-based World Models? Paper • 2512.18832 • Published Dec 21, 2025 • 16 • 3
InCoder-32B: Code Foundation Model for Industrial Scenarios Paper • 2603.16790 • Published Mar 17 • 312