DeepSeek-V4.1-Flash: Pushing the Limits of KV Cache Compression Paper • 2609.19969 • Published 4 days ago • 134
SenseNova-U1.5: Towards Native Unified Visual Intelligence Paper • 2609.11929 • Published 11 days ago • 270
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published 20 days ago • 24
Code as Worlds: Agentic Discovery of Executable World Representations for Physical Reasoning Paper • 2608.27549 • Published 25 days ago • 49
Qwen-UI-Agent Technical Report: Toward Next-Generation Real-World Centric Foundation GUI Agents Paper • 2607.28227 • Published Jul 30 • 157
Harness Handbook: Making Evolving Agent Harnesses Readable,Navigable, and Editable Paper • 2607.13285 • Published Jul 14 • 125
OpenVisTool: An Open Recipe for Synthesizing Instructive Visual Tool-Use Trajectories Paper • 2608.08557 • Published Aug 9 • 2
Enjoy Your Talk: A Human-Centered Benchmark for Multi-Turn Dialogue with Decoupled User Simulation, Target Modeling, and Judging Paper • 2607.10428 • Published Jul 24 • 1
Omni-Perception Policy Optimization for Multimodal Emotion Reasoning Paper • 2606.25325 • Published Jun 24
MER-R1: Multimodal Emotion Reasoning via Slow-Fast Thinking Synergy Paper • 2606.27652 • Published Jun 26
EgoPro-Bench: Benchmarking Personalized Proactive Interaction in Egocentric Video Streams Paper • 2605.07299 • Published May 8
ISE: An Execution-Grounded Recipe for Multi-Turn OS-Agent Trajectories Paper • 2606.11520 • Published Jun 9
From Pixels to Words -- Towards Native One-Vision Models at Scale Paper • 2605.28820 • Published May 27 • 75