OmniConfess: Eliciting Token Confessions to Mitigate Omni-Modal Hallucination Paper • 2610.02999 • Published 5 days ago • 7
HelixWorld: A Real-time Interactive Audio-Visual World Model Paper • 2609.38123 • Published 11 days ago • 37
ActionPiece: Rethinking Action Tokenization for Autoregressive Vision-Language-Action Models Paper • 2609.18487 • Published 24 days ago • 48
Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding Paper • 2609.04131 • Published Sep 3 • 31
SemComp-Bench: Benchmarking Semantic Task Completion in Video Generation Paper • 2608.17426 • Published Aug 18 • 161
Can We Defend Against AI-Generated Video Attacks on Real-World Crisis Events? A Systematic Evaluation of Detectors, Generators and Social Dissemination Paper • 2608.14391 • Published Aug 14 • 287
UniSwap: Streaming Audio-Visual Identity Swapping for Talking Videos Paper • 2608.11752 • Published Aug 13 • 23
Context-Matched Distillation: Teacher Causality for Autoregressive Video Distillation Paper • 2608.13391 • Published Aug 13 • 20
OpenART: Scaling Agent Red Teaming via Open-Ended Environment Evolution Paper • 2608.00677 • Published Aug 1 • 266
JigShape: Evaluating Visual-Geometric Reasoning in VLMs through Jigsaw Puzzles Paper • 2607.27670 • Published Aug 4 • 9
SymDiag: Explainable Diagnosis for LLM Reasoning via Neuro-Symbolic Verification Paper • 2608.08786 • Published Aug 9 • 6
SKT: Skill-Use Training at Scale via Verified Synthetic Data Generation Paper • 2608.02287 • Published Aug 3 • 32
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing Paper • 2607.29025 • Published Jul 31 • 18