FastVideo/FastVideo-FastH3-4-step-Preview-v1-VSA-DataFree Text-to-Video • 35B • Updated 30 days ago • 1.11M • 318
view article Article Efficient Request Queueing – Optimizing LLM Performance tngtech • Apr 2, 2025 • 29
view article Article Prefill and Decode for Concurrent Requests - Optimizing LLM Performance tngtech • Apr 16, 2025 • 99
view article Article A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes ybelkada, timdettmers • Aug 17, 2022 • 140
view article Article Making LLMs even more accessible with bitsandbytes, 4-bit quantization and QLoRA +3 ybelkada, timdettmers, artidoro, sgugger, smangrul • May 24, 2023 • 182
view article Article Mixture of Experts Explained +4 osanseviero, lewtun, philschmid, smangrul, ybelkada, pcuenq • Dec 11, 2023 • 1.19k