-
MACRO: Advancing Multi-Reference Image Generation with Structured Long-Context Data
Paper • 2603.25319 • Published • 32 -
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale
Paper • 2603.25040 • Published • 124 -
MinerU-Diffusion: Rethinking Document OCR as Inverse Rendering via Diffusion Decoding
Paper • 2603.22458 • Published • 131 -
Speed by Simplicity: A Single-Stream Architecture for Fast Audio-Video Generative Foundation Model
Paper • 2603.21986 • Published • 120
Cai
Mona8834
AI & ML interests
None yet
Recent Activity
updated a collection 3 days ago
LLM Papers updated a collection 7 days ago
LLM Papers updated a collection 7 days ago
LLM PapersOrganizations
None yet