view article Article **Know Who Spoke When: Build Real-Time, Multi-Speaker AI with NVIDIA Nemotron 3 Diarization** nvidia • 2 days ago • 51
view article Article tokenizers v1: encode, decode and scaling, measured +2 ArthurZ, sbrandeis, mcpotato, lysandre • 5 days ago • 74
view article Article Transformers now runs llama.cpp quants +1 marcsun13, ArthurZ, lysandre • 4 days ago • 62
view article Article IBM releases SOTA Granite Time Series PatchTST-FM-r2 model with commercial-friendly license ibm-research • 16 days ago • 60
view article Article Wire It, Run It, Deploy It: AI Workflows in Gradio ysharma, abidlabs • Aug 25 • 49
view article Article Training and Finetuning Multi-Vector Embedding Models with Sentence Transformers tomaarsen • about 1 month ago • 144
view article Article Introducing @huggingface/kernels: 200+ WebGPU Kernels for Local AI nico-martin, Xenova • 25 days ago • 82
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 23 days ago • 135
VideoChat3: Fully Open Video MLLM for Efficient and Generalist Video Understanding Paper • 2607.14935 • Published Jul 16 • 120
view article Article Multimodal Embedding & Reranker Models with Sentence Transformers tomaarsen • Apr 9 • 75
video-SALMONN 2 Collection video-SALMONN 2 is a powerful audio-visual large language model (LLM) that generates high-quality audio-visual video captions. • 11 items • Updated Mar 21 • 2
view article Article VLX-Flow: Continuous Video Understanding for Real-Time Multimodal Interaction omlab • Jun 27 • 15
view article Article Profiling in PyTorch (Part 3): Attention is all you profile +2 ariG23498, sergiopaniego, sayakpaul, ror • Jul 10 • 50
Granite 4.1 Language Models Collection Efficient language models for multilingual generation, coding, RAG, and AI assistant workflows. • 15 items • Updated 11 days ago • 70
ReaderLM-v2: Small Language Model for HTML to Markdown and JSON Paper • 2503.01151 • Published Mar 3, 2025 • 6