StableVQ: Practical Guidelines for Stable Vector-Quantized Tokenizer Training Paper • 2609.26774 • Published 10 days ago • 56
CodeMidas: Scaling Agentic Coding RL Environments from Code Itself Paper • 2609.22068 • Published 14 days ago • 138
Learning from the Self-future: On-policy Self-distillation for dLLMs Paper • 2606.18195 • Published Jun 16 • 42
Rapidata/2k-ranked-images-open-image-preferences-v1 Viewer • Updated Apr 10, 2025 • 2k • 34 • 27
Intern-S1-Pro: Scientific Multimodal Foundation Model at Trillion Scale Paper • 2603.25040 • Published Mar 26 • 132
Stepping VLMs onto the Court: Benchmarking Spatial Intelligence in Sports Paper • 2603.09896 • Published Mar 10 • 28
MOVA: Towards Scalable and Synchronized Video-Audio Generation Paper • 2602.08794 • Published Feb 9 • 159
Render-of-Thought: Rendering Textual Chain-of-Thought as Images for Visual Latent Reasoning Paper • 2601.14750 • Published Jan 21 • 18
Boosting Latent Diffusion Models via Disentangled Representation Alignment Paper • 2601.05823 • Published Jan 9 • 17
RecTok: Reconstruction Distillation along Rectified Flow Paper • 2512.13421 • Published Dec 15, 2025 • 5
SVG-T2I: Scaling Up Text-to-Image Latent Diffusion Model Without Variational Autoencoder Paper • 2512.11749 • Published Dec 12, 2025 • 39
VQRAE: Representation Quantization Autoencoders for Multimodal Understanding, Generation and Reconstruction Paper • 2511.23386 • Published Nov 28, 2025 • 16
A Style is Worth One Code: Unlocking Code-to-Style Image Generation with Discrete Style Space Paper • 2511.10555 • Published Nov 13, 2025 • 64
Group Relative Attention Guidance for Image Editing Paper • 2510.24657 • Published Oct 28, 2025 • 26
Sample By Step, Optimize By Chunk: Chunk-Level GRPO For Text-to-Image Generation Paper • 2510.21583 • Published Oct 24, 2025 • 31