view article Article **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** nvidia • 6 days ago • 63
onPanda: Efficient Annotation of On-Policy Alignment Data for LLMs and Agents via Token-Level Correction Paper • 2609.24983 • Published 8 days ago • 55
Unlocking Lossless Speedups in LLMs via Discrete Diffusion Paper • 2609.04010 • Published 26 days ago • 114
TerraMind: Large-Scale Generative Multimodality for Earth Observation Paper • 2504.11171 • Published Apr 15, 2025 • 3
AuK Technical Report: An Open-Source Foundational Model for Speech Generation and Editing Paper • 2609.08936 • Published 21 days ago • 165
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 26 days ago • 139
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published 26 days ago • 186
K2 Horizon Collection K2 Horizon models, datasets, and supporting resources • 22 items • Updated 18 days ago • 137
H3-World: Turning Language Understanding into World Control Paper • 2609.01560 • Published 28 days ago • 52
LAION-BVD: A 10-Million-Hour Open Video Dataset for Multimodal Pre-training Paper • 2608.24845 • Published Aug 25 • 16
WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report Paper • 2608.24053 • Published Aug 25 • 71
Granite 4.2 Language Models Collection Efficient reasoning and thinking language models for multilingual generation, coding, and AI assistant workflows. • 24 items • Updated 14 days ago • 43