UniWorld-Design: From Pixel Generation to Layer-Native Design Paper • 2608.03971 • Published 3 days ago • 19
UniWorld-Design: From Pixel Generation to Layer-Native Design Paper • 2608.03971 • Published 3 days ago • 19
RareLens: Towards End-to-End Rare Disease Care via Aligning Divergent Large Language Model Reasoning Paper • 2607.23290 • Published 13 days ago • 1
Agentifying Patient Dynamics within LLMs through Interacting with Clinical World Model Paper • 2605.14723 • Published May 14
GameCraft-Bench: Can Agents Build Playable Games End-to-End in a Real Game Engine? Paper • 2606.17861 • Published Jun 16 • 59
WaveMind: Towards a Conversational EEG Foundation Model Aligned to Textual and Visual Modalities Paper • 2510.00032 • Published Sep 26, 2025
MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation Paper • 2603.00585 • Published Feb 28 • 2
MicroVerse: A Preliminary Exploration Toward a Micro-World Simulation Paper • 2603.00585 • Published Feb 28 • 2
iFSQ: Improving FSQ for Image Generation with 1 Line of Code Paper • 2601.17124 • Published Jan 23 • 34
Does Understanding Inform Generation in Unified Multimodal Models? From Analysis to Path Forward Paper • 2511.20561 • Published Nov 25, 2025 • 33
Uniworld-V2: Reinforce Image Editing with Diffusion Negative-aware Finetuning and MLLM Implicit Feedback Paper • 2510.16888 • Published Oct 19, 2025 • 22
Can Understanding and Generation Truly Benefit Together -- or Just Coexist? Paper • 2509.09666 • Published Sep 11, 2025 • 34
ImgEdit: A Unified Image Editing Dataset and Benchmark Paper • 2505.20275 • Published May 26, 2025 • 20
ShizhenGPT: Towards Multimodal LLMs for Traditional Chinese Medicine Paper • 2508.14706 • Published Aug 20, 2025
TalkVid: A Large-Scale Diversified Dataset for Audio-Driven Talking Head Synthesis Paper • 2508.13618 • Published Aug 19, 2025 • 19
MedGen: Unlocking Medical Video Generation by Scaling Granularly-annotated Medical Videos Paper • 2507.05675 • Published Jul 8, 2025 • 27
UniWorld: High-Resolution Semantic Encoders for Unified Visual Understanding and Generation Paper • 2506.03147 • Published Jun 3, 2025 • 59
UPME: An Unsupervised Peer Review Framework for Multimodal Large Language Model Evaluation Paper • 2503.14941 • Published Mar 19, 2025 • 5
WISE: A World Knowledge-Informed Semantic Evaluation for Text-to-Image Generation Paper • 2503.07265 • Published Mar 10, 2025 • 4