ToolArtist: Tool-Using Unified Multimodal Models for Agentic Image Generation Paper • 2608.04436 • Published 8 days ago • 58
JoyAI-Video-Edit: Real-Time Open-Ended Video Editing with Autoregressive Diffusion Paper • 2608.03974 • Published 9 days ago • 91
PosterMELD: Multi-Agent Paper-to-Poster Generation for Controllable Design Diversity with Editable Print-Ready Outputs Paper • 2608.02218 • Published 10 days ago • 7
Hunyuan3D-Buffalo 1.0: A Unified Multimodal Model for Scalable 3D Generation, Understanding, and Editing Paper • 2608.02711 • Published 9 days ago • 88
UniWorld-Design: From Pixel Generation to Layer-Native Design Paper • 2608.03971 • Published 8 days ago • 21
AURORA-LM: Autoencoding Unified Representation for Continuous-Latent Diffusion Language Modeling Paper • 2608.02602 • Published 10 days ago • 79
CAPEval: A Decoupled Caption Evaluation across Understanding and Generation Paper • 2608.02589 • Published 10 days ago • 25
Any-OPD: Heterogeneous On-Policy Distillation for Flow-Matching Models via Representation-Space Bridging Paper • 2608.03316 • Published 9 days ago • 26
SwanTale: Unified Multi-Speaker Speech and Audio Generation for Instruct and Zero-Shot Tasks Paper • 2608.02023 • Published 10 days ago • 155
Evaluation-Verification Reward for Consistent Multi-Reference Image Editing Paper • 2607.29025 • Published 13 days ago • 16
Frontis-MA1: Training an AI4AI Model towards Recursive Self-Improvement in Machine Learning Engineering Paper • 2607.28568 • Published 14 days ago • 182
HumanCLAW: Can Vision-Language Models Act Through a Body? Paper • 2607.27180 • Published 15 days ago • 76
HiFi-UMI: Learning Deployable Manipulation Policies from High-Fidelity UMI Data Alone Paper • 2607.25895 • Published 16 days ago • 156
TILT: Improving Compositional Generation in Diffusion Models with a Model-Intrinsic Reward Paper • 2607.21606 • Published May 16 • 3
Spectral Prior for Reducing Exposure Bias in Diffusion Models Paper • 2607.22091 • Published 20 days ago • 6
Skill Self-Play: Pushing the Frontier of LLM Capability with Co-Evolving Skills Paper • 2607.22529 • Published 20 days ago • 48
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment Paper • 2607.13429 • Published 29 days ago • 16