view article Article Native-speed vLLM transformers modeling backend hmellor, lysandre • 14 days ago • 56
view article Article Harness, Scaffold, and the AI Agent Terms Worth Getting Right sergiopaniego, ariG23498 • May 25 • 134
view article Article PhysicsIntern: from an Autonomous Benchmark-runner to a Research Sidekick dlouapre • Jun 11 • 7
view article Article Shipping a Trillion Parameters With a Hub Bucket: Delta Weight Sync in TRL +6 aminediroHF, qgallouedec, kashif, lewtun, edbeeching, albertvillanova, lvwerra, sergiopaniego • May 27 • 43
🧬 Carbon Collection Carbon 500M, 3B, 8B genomic models and GGUF variants for llama.cpp • 7 items • Updated Jun 2 • 44
view article Article Two Years of Local AI on a Laptop: When Open Models Outpaced Moore's Law mishig • May 11 • 24
Scaling LLM Test-Time Compute Optimally can be More Effective than Scaling Model Parameters Paper • 2408.03314 • Published Aug 6, 2024 • 67
Embarrassingly Simple Self-Distillation Improves Code Generation Paper • 2604.01193 • Published Apr 1 • 56
view article Article How we OCR'ed 30,000 papers using Codex, open OCR models and Jobs nielsr • Apr 7 • 62
view article Article How I contributed a new model to the Transformers library using Codex nielsr • Mar 30 • 52