Frontier Language Models Struggle to Copy: Text Can Be Better Viewed in 2D Paper • 2607.16072 • Published 25 days ago • 1
Bad Memory: Evaluating Prompt Injection Risks from Memory in Agentic Systems Paper • 2607.14611 • Published 26 days ago • 1
PAST-Bench: Benchmarking the Foundations of Recursive Self-Improvement in Personal Agents Paper • 2608.04003 • Published 7 days ago • 33
When Attention Goes Blind: Numerical Failure in ALiBi Positional Encodings Paper • 2608.03994 • Published 7 days ago • 8
Distilled Reinforcement Learning for LLM Post-training Paper • 2607.17247 • Published 23 days ago • 11
LLM-as-a-Coach: Experiential Learning for Non-Verifiable Tasks Paper • 2607.18110 • Published 22 days ago • 16
NVIDIA-labs OO Agents: Native Python Object-Oriented Agents Paper • 2607.20709 • Published 20 days ago • 34
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published 6 days ago • 58
ContinualSkillBench: Can LLM Agents Truly Evolve Their Capabilities? Paper • 2608.03874 • Published 7 days ago • 14
Permission Denied: Policy-Graded Evaluation of Coding Agents in Hardened Environments Paper • 2608.02670 • Published 9 days ago • 1
DataSpace: Benchmarking Data Agents for Verifiable Analytics over Heterogeneous Workspaces Paper • 2608.03451 • Published 7 days ago • 30
Harness-R1: Learning to Edit Executable Runtime Harnesses from Agent Failure Trajectories Paper • 2608.02276 • Published 8 days ago • 3
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Paper • 2607.23802 • Published 16 days ago • 105
Partition, Prompt, Aggregate: Statistical Self-Consistency in Language Models Paper • 2607.15277 • Published 26 days ago • 10
MetaView: Monocular Novel View Synthesis with Scale-Aware Implicit Geometry Priors Paper • 2607.12000 • Published 29 days ago • 40
LATO.2: Factorized 3D Mesh Generation with Vertex and Topology Flow Paper • 2607.10623 • Published about 1 month ago • 16
Hierarchical Denoising For Multi-Step Visual Reasoning Paper • 2607.15278 • Published 26 days ago • 7
From Pixels to States: Rethinking Interactive World Models as Game Engines Paper • 2607.14076 • Published 27 days ago • 37
BadWAM: When World-Action Models Dream Right but Act Wrong Paper • 2607.15207 • Published 26 days ago • 54