Rethinking Token Reweighting for SFT: Suppress, Reverse, and Extrapolate Learned Features Paper • 2609.33463 • Published 9 days ago • 11
AutoGUIWorld: Image Generators as Visual World Models for GUI Agent Paper • 2610.01215 • Published 5 days ago • 57
VABench: Measuring Embodied Spatial Intelligence through Visual Demonstrations, Active Perception, and Metric Control Paper • 2609.19554 • Published 19 days ago • 43
Generative Late-Interaction Embeddings For Visual Document Retrieval Paper • 2609.11808 • Published 26 days ago • 27
Cadence: Error-Bounded Lossy Compression of Demand Time Series with a Time-Series Foundation Model Paper • 2609.06008 • Published Sep 5 • 19
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Training Chemical Plausibility-Aware Large Language Models for Single-Step Retrosynthesis Paper • 2608.18940 • Published Aug 19 • 35
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
PlayWorld: Benchmarking World Models with Agent Players over Long-Horizon Objectives Paper • 2608.13552 • Published Aug 13 • 47
LycheeMemory V2: Efficient Long-Term Memory for LLM Agents via Semantic Segment-Level Consolidation Paper • 2608.12990 • Published Aug 13 • 14
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
AdvFD: Boosting Visual Generation via Adversarial Fr'echet Distance Loss Paper • 2608.11205 • Published Aug 11 • 27
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning Paper • 2608.06197 • Published Aug 6 • 48
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 107
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published Aug 3 • 161