A2Z GameSpec-Bench: How Faithfully Can Coding Agents Generate Games from Game Design Specifications? Paper • 2609.39564 • Published 11 days ago • 18
PivotOPD: Learning to Recover from Pivotal Mistakes in Multi-Turn Agents Paper • 2609.40285 • Published 11 days ago • 23
Agent Error Dataset: Scaling 50,000 Error--Diagnosis Pairs for Failure Analysis and Error-Aware Post-Training Paper • 2609.40111 • Published 11 days ago • 51
Thinking Outside the Box: Can Language Models Rely on External Guidance Selectively? Paper • 2609.39578 • Published 11 days ago • 72
RSIGame: Autonomous Agentic Game Development with Recursive Self-improvement Paper • 2609.39045 • Published 11 days ago • 93
Rethinking Latent Visual Reasoning: Grounding Latent Reasoning in Visual Evidence Paper • 2609.34563 • Published 13 days ago • 219