InternW0-Δ: A World Action Model Bridging Predictive Dynamics and Actions with 20K+ Hours of Open Data Paper • 2609.31394 • Published 7 days ago • 27
Game Arena: Strategic LLM Evaluation in Competitive Environments Paper • 2609.31473 • Published 7 days ago • 12
Uncovering Understanding-Generation Synergy in Native Unified Multimodal Models: From Representation, Task to System Paper • 2609.01607 • Published about 1 month ago • 24
On the Design Fundamentals of Pixel Text Representation Learning Paper • 2609.01147 • Published about 1 month ago • 31
FACET: Preserving Source Intent and Executable State in Terminal Task Synthesis Paper • 2608.18580 • Published Aug 19 • 96
An Empirical Study of Training Pixel-Space Text-to-Image Diffusion Models Paper • 2608.16887 • Published Aug 17 • 36
Towards Physics of Multimodal Pretraining: Knowledge Flow, Modality Synergy, Early Unification, and Recipes Paper • 2608.05000 • Published Aug 6 • 65
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published Aug 4 • 26
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published Aug 4 • 107
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published Aug 3 • 85
AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities Paper • 2607.13705 • Published Jul 15 • 40
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Paper • 2607.07708 • Published Jul 8 • 88
Scaling Mixture-of-Experts Video Pretraining for Embodied Intelligence Paper • 2607.07675 • Published Jul 8 • 64