WorldLine: Action-Driven Visual Simulation for Robotic Manipulation Paper • 2609.38059 • Published 4 days ago • 24
MaLiang-Harness: A Programmable Path to Image and Video Generation Paper • 2609.34309 • Published 5 days ago • 400
SoL-Refiner: Speed-of-Light One-Step Refinement for High-Resolution Video Paper • 2609.37969 • Published 4 days ago • 34
LongLive-Plug: Once-for-All Distillation for Video Generation Paper • 2609.38154 • Published 4 days ago • 34
VoxMem: Benchmarking Multimodal Memory in Large Audio Language Models Paper • 2609.32607 • Published 7 days ago • 144
VideoPhysEdit: Physical Counterfactual Video Editing via Rigid-Body Physical Scene Reconstruction Paper • 2609.35134 • Published 5 days ago • 19
SpatialSpeak: QA-Native Reconstruction with Local and Global Context for Spatial Chain-of-Thought Reasoning Paper • 2609.33616 • Published 6 days ago • 25
TimeEvo: Failure-Driven Self-Evolution of a Time Series Agent Paper • 2609.27277 • Published 10 days ago • 32
ViRDM: Taming Representation Distribution Matching for Few-Step Causal Video Generation Paper • 2609.28923 • Published 9 days ago • 10
AV-GRPO: Modality-Anchored Decoupling Diffusion Reinforcement Learning for Joint Audio-Video Generation Paper • 2609.29816 • Published 9 days ago • 12
WanPE: Towards Cinematic Prompt Enhancement for Modern Text-to-Video Generation Paper • 2609.30221 • Published 9 days ago • 46
RewardVerse: Rubric-Guided Policy Optimization for Video Reward Modeling Paper • 2609.22947 • Published 14 days ago • 42
All modalities are equal, but video is more equal: Closing the Cross-Attention Gap in Joint Video Generation Paper • 2609.27901 • Published 10 days ago • 22
The Past Frames the Future: Memory for Autoregressive Video Generation Paper • 2609.28466 • Published 10 days ago • 64
MemoryAthena: Adaptive Routing over Latent and Generated Memories Paper • 2609.25853 • Published 11 days ago • 12
OmniVBench: A Benchmark and Large-Scale Dataset for Omni Reference-to-Video Generation Paper • 2609.22069 • Published 15 days ago • 37
VC-Attention: Value Smoothing and Softmax Casting for Low-bit Attention Paper • 2609.15810 • Published 19 days ago • 51
OmniVChat: Synthesizing, Benchmarking, and Training for Native Audio-Visual Dialogue Paper • 2609.21465 • Published 15 days ago • 150