Hiding Tool Latency in On-Device Cascaded Voice Agent through Speculative Execution Paper • 2610.07641 • Published 5 days ago • 14
Rethinking Cross-Tokenizer On-Policy Distillation: From Alignment Coverage to Supervision Reliability Paper • 2610.08448 • Published 5 days ago • 200
Prism: Dynamic Sparse Attention for Native 2K Joint Video-Audio Generation Model Training Paper • 2610.05416 • Published 7 days ago • 20
Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation Paper • 2610.05608 • Published 7 days ago • 162
Representation-Space MMD for Diffusion Language Models Paper • 2610.06648 • Published 6 days ago • 31
Latent-MOPD: Latent Multi-Teacher On-Policy Distillation Paper • 2610.02381 • Published 10 days ago • 70
Unmask the State: When Does State Adaptation Matter for Masked Diffusion Language Models Paper • 2609.33355 • Published 14 days ago • 51
Scheduling Recursive Reasoning in Looped Transformers Paper • 2609.36653 • Published 12 days ago • 20
Beyond the Timeline: Augmenting Long-Video Memory with Grounded Entity Biographies Paper • 2609.38155 • Published 12 days ago • 115
CrossBFM: Distilling a Shared Latent Behavior Space Across Humanoid Embodiments Paper • 2609.38087 • Published 12 days ago • 25
Skill2Env: Capability-Oriented Environment Synthesis from Skills for General Agents Paper • 2609.33772 • Published 14 days ago • 35
Enhancing Photogrammetric Digital Surface Models with Pretrained Diffusion Models and Multimodal Conditioning Paper • 2609.31199 • Published 16 days ago • 14
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 18 days ago • 24
The Past Frames the Future: Memory for Autoregressive Video Generation Paper • 2609.28466 • Published 18 days ago • 65
HuRo: Robotizing Human Videos for Scalable VLA Pretraining Paper • 2609.10706 • Published 23 days ago • 29
An Empirical Study of Harness Design for Coding Agents Paper • 2609.20804 • Published 24 days ago • 94
Ovis-Embedding: Pushing the Frontiers of Universal Omni-Modal Embeddings Paper • 2609.25165 • Published 20 days ago • 77
WorldCrafter: Consistent Video World Model with Implicit 3D-aware Memory Paper • 2609.24984 • Published 20 days ago • 158
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 23 days ago • 153