Multimodal Flow: Unified Flow Modeling of Language and Vision in Embedding Spaces Paper • 2609.40362 • Published 3 days ago • 14
SemanTok: Predictable Semantic Tokens for Efficient Autoregressive Video Generation Paper • 2610.00686 • Published 3 days ago • 4
view article Article FLUX 3 Action: a world action model you can fine-tune black-forest-labs • 9 days ago • 13
view article Article Open TTS Leaderboard: Scalable Evaluation for Multilingual Text-to-Speech and Voice Cloning +2 bezzam, Steveeeeeeen, eustlb, mrfakename • 3 days ago • 36
TT-VidT: Decoupling the Temporal Axis for Efficient Motion-Centric Video Pretraining Paper • 2609.33419 • Published 6 days ago • 18
UniMate Collection UniMate: One Unified Model to Animate Diverse Skeletons (SIGGRAPH Asia 2026). Paper & project page: https://linzhanmou.com/unimate/ • 7 items • Updated 5 days ago • 15
InternW0-Δ Collection The collection includes the pretrained and finetuned checkpoints of InternW0-Δ • 4 items • Updated 5 days ago • 4
view article Article YODAS v3: A 1 Million Hour Dataset for the Next Generation of Open Voice AI Research espnet • 5 days ago • 30
Ming-Image (MLX) Collection Ming-Image-0.1 (MIT): design text-to-image with native RGBA + layer decomposition. Swift/MLX: github.com/xocialize/ming-image-swift • 6 items • Updated 5 days ago • 1
WISA: World Simulator Assistant for Physics-Aware Text-to-Video Generation Paper • 2503.08153 • Published Mar 11, 2025 • 4
SmolDataEnvs Collection 5.5K+ RL tasks for hill-climbing small models in code and data science. Deterministic grading, no LLM judge. • 10 items • Updated 1 day ago • 8
RoboChallenge: Large-scale Real-robot Evaluation of Embodied Policies Paper • 2510.17950 • Published Oct 20, 2025 • 11
DDM Actions Artifacts Collection Canonical DDM Actions models and caches mapped to GitHub families. • 4 items • Updated 6 days ago • 1
Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms Paper • 2609.23658 • Published 13 days ago • 30
FLUX 3 Action Collection Open weights 7B world action model: base and shared encoders, SO-101 and DROID policies. Code: https://github.com/black-forest-labs/flux-action • 3 items • Updated about 9 hours ago • 41
view article Article tokenizers v1: encode, decode and scaling, measured +2 ArthurZ, sbrandeis, mcpotato, lysandre • 12 days ago • 86