DecepEval: A Benchmark for Evaluating Deception in LLM Agents Paper • 2610.07967 • Published 3 days ago • 67
ScholarCatalyst: A Benchmark for Retrieving Papers That Inspire New Research Paper • 2610.02202 • Published 8 days ago • 13
Confidence Comes from Experience: Experiential Confidence Estimation from Reasoning to Agents Paper • 2609.17708 • Published 24 days ago • 78
Why Gated DeltaNet Survives 4-Bit Quantization: NVFP4 W4A4 for the Recurrent Half of a Hybrid 27B LLM Paper • 2609.04098 • Published Sep 3 • 86
Repo-To-Skill: Distilling GitHub Repositories Into AI4AI Skills Paper • 2609.02749 • Published Sep 2 • 408
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
ChronoVision: Temporal Reasoning via Latent State Reconstruction Paper • 2608.05631 • Published Aug 6 • 40
When Memory Lies: An Empirical Study of Spatial Memory Staleness in VLM Agents Paper • 2608.04574 • Published Aug 5 • 16
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published Aug 3 • 85