DecepEval: A Benchmark for Evaluating Deception in LLM Agents Paper • 2610.07967 • Published 3 days ago • 67
OpenTumorBoard: A Real-World Benchmark of Multidisciplinary Tumor Board Discussion Trajectories Paper • 2609.32810 • Published 13 days ago • 15
SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem Paper • 2609.07064 • Published Sep 7 • 147
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published Sep 3 • 86
CoToGrasp: Contact-Topology-Conditioned Dexterous Grasp Synthesis via Canonical Workspace Learning Paper • 2608.19776 • Published Aug 20 • 7
Massive Activations in Hybrid Linear Attention Large Language Models: Pre-Attention Spikes and Inter-Spike Plateaus Paper • 2608.12149 • Published Aug 12 • 30
On-Policy Delta Distillation for Multilingual Math Reasoning Paper • 2608.05802 • Published Aug 6 • 33
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation Paper • 2608.06374 • Published Aug 6 • 23
Scaling Properties of Text Conditioning in Visual Generation Paper • 2607.29679 • Published Jul 31 • 42