PackLab: A Comprehensive Framework for Developing, Training, and Evaluating MLLMs in Robotic Bin Packing Paper • 2609.23784 • Published 8 days ago • 11
SoL-Pi: Recursively Scaling Auto-Research Loops for Efficient Agent Harness Paper • 2609.20519 • Published 11 days ago • 136
Annotations as Rollouts: Efficient and Scalable Reinforcement Learning for Video MLLMs Paper • 2608.20492 • Published Aug 20 • 87
DeepSearch-World: Self-Distillation for Deep Search Agents in a Verifiable Environment Paper • 2607.07820 • Published Jul 8 • 95
Accurate, Interdisciplinary and Transparent Structure-property Understanding with Deep Native Structural Reasoning Paper • 2607.07708 • Published Jul 8 • 88
SANA-Streaming: Real-time Streaming Video Editing with Hybrid Diffusion Transformer Paper • 2605.30409 • Published May 28 • 39
SkillEvolBench: Benchmarking the Evolution from Episodic Experience to Procedural Skills Paper • 2605.24117 • Published May 22 • 20
GenEvolve: Self-Evolving Image Generation Agents via Tool-Orchestrated Visual Experience Distillation Paper • 2605.21605 • Published May 20 • 15
SANA-WM: Efficient Minute-Scale World Modeling with Hybrid Linear Diffusion Transformer Paper • 2605.15178 • Published May 14 • 92
OmniShow: Unifying Multimodal Conditions for Human-Object Interaction Video Generation Paper • 2604.11804 • Published Apr 13 • 73
PosterIQ: A Design Perspective Benchmark for Poster Understanding and Generation Paper • 2603.24078 • Published Mar 25 • 1
RoomPlanner: Explicit Layout Planner for Easier LLM-Driven 3D Room Generation Paper • 2511.17048 • Published Nov 21, 2025 • 2
UniFit: Towards Universal Virtual Try-on with MLLM-Guided Semantic Alignment Paper • 2511.15831 • Published Nov 19, 2025 • 2
Garments2Look: A Multi-Reference Dataset for High-Fidelity Outfit-Level Virtual Try-On with Clothing and Accessories Paper • 2603.14153 • Published Mar 14 • 3
Emotion-Director: Bridging Affective Shortcut in Emotion-Oriented Image Generation Paper • 2512.19479 • Published Dec 22, 2025 • 1
HiFi-Inpaint: Towards High-Fidelity Reference-Based Inpainting for Generating Detail-Preserving Human-Product Images Paper • 2603.02210 • Published Mar 2 • 30
Modality Gap-Driven Subspace Alignment Training Paradigm For Multimodal Large Language Models Paper • 2602.07026 • Published Feb 2 • 141