Can MiniMax-H3 Reason About the Physical World? An Evaluation of Omni-Modal Generative Model Paper • 2609.18323 • Published 10 days ago • 130
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published 25 days ago • 52
Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence Paper • 2608.31075 • Published 26 days ago • 29
Z-Erase: Enabling Concept Erasure in Single-Stream Diffusion Transformers Paper • 2603.25074 • Published May 10
WorldMind: Decoupled Game World Model for State-Aware NPC Behavior Paper • 2608.21439 • Published Aug 18 • 15
StatePlay: State-Aware Game World Models for Mechanics-Consistent Generation Paper • 2607.26754 • Published Jul 29 • 19
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published Jul 27 • 39
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published Jul 26 • 127
Sol-Attn: Accelerating Video Generation Inference via On-the-Fly Attention Sparsification Paper • 2607.24027 • Published Jul 27 • 39
JarvisHub: An Open Harness for Canvas-Native Multimodal Creative Agents Paper • 2607.23588 • Published Jul 26 • 127
GEAR: Guided End-to-End AutoRegression for Image Synthesis Paper • 2606.32039 • Published Jun 30 • 34
SCOPE: Simulating Cross-game Operations in Playable Environments for FPS World Models Paper • 2605.23345 • Published May 22 • 16
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation Paper • 2605.21605 • Published May 20 • 15
Mem4D: Decoupling Static and Dynamic Memory for Dynamic Scene Reconstruction Paper • 2508.07908 • Published Aug 11, 2025
Fast or Slow? Integrating Fast Intuition and Deliberate Thinking for Enhancing Visual Question Answering Paper • 2506.00806 • Published Jun 1, 2025
UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment Paper • 2511.15831 • Published Nov 19, 2025 • 2
Incantation: Natural Language as the Action Interface for Multi-Entity Video World Models Paper • 2605.18601 • Published May 18 • 6