RealSWE: A Compositional Evaluation of Coding Agents under Realistic User Requests Paper • 2608.27831 • Published 26 days ago • 33
view article Article Extremely Fast and Accurate Transcription with Granite Speech 5.0 Turbo CTC ibm-granite • Aug 25 • 36
Agent Memory Distillation: Empowering Small LLM Agents with Hierarchical Teacher Memory Paper • 2608.07169 • Published Aug 7 • 51
StreamArena: Toward Continuous, Interactive, and Long-Horizon Agentic Streaming Video Understanding Paper • 2608.05703 • Published Aug 6 • 17
Beyond Borrowed Histories: Person-Aligned User Simulation for Interactive Role-Playing Evaluation Paper • 2607.27816 • Published Jul 30 • 35
BM25 Wins at Scale: A Scaling Study of Retrieval-Augmented Generation Paradigms Paper • 2607.26497 • Published Jul 30 • 53
EvolvingWorld: An Open-Schema Framework for Co-Evolving Role-Play Agents and World Model in Interactive Literary World Paper • 2607.17250 • Published Jul 19 • 93
Reinforcement Learning with Metacognitive Feedback Elicits Faithful Uncertainty Expression in LLMs Paper • 2606.32032 • Published Jun 30 • 29
Adam's Law: Textual Frequency Law on Large Language Models Paper • 2604.02176 • Published Apr 2 • 110
ClawGUI: A Unified Framework for Training, Evaluating, and Deploying GUI Agents Paper • 2604.11784 • Published Apr 13 • 143
GameWorld: Towards Standardized and Verifiable Evaluation of Multimodal Game Agents Paper • 2604.07429 • Published Apr 8 • 63
Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid Mamba-Transformer Model for Agentic Reasoning Paper • 2604.12374 • Published Apr 14 • 39
Learn2Fold: Structured Origami Generation with World Model Planning Paper • 2603.29585 • Published Feb 2 • 16
XSkill: Continual Learning from Experience and Skills in Multimodal Agents Paper • 2603.12056 • Published Jul 1 • 34