view article Article Continuous batching from first principles +1 ror, ArthurZ, mcpotato • Nov 25, 2025 • 444
Distilling an End-to-End Voice Assistant Without Instruction Training Data Paper • 2410.02678 • Published Oct 3, 2024 • 24
Molmo Collection Artifacts for open multimodal language models. • 5 items • Updated Dec 23, 2025 • 310
view article Article The case for specialized pre-training: ultra-fast foundation models for dedicated tasks Pclanglais • Aug 4, 2024 • 30
sentence-transformers-from-synthetic-data Collection Example of using distilabel to generate synthetic triplets data for fine-tuning a Sentence Transformer model • 4 items • Updated Jun 21, 2024 • 23