Grouped Value Attention: Efficient KV Caching via On-Demand Key Reconstruction Paper • 2609.13285 • Published 17 days ago • 81
FrontiersMind/Nandi-Mini-V1.1-600M-Early-Checkpoint-250GT Text Generation • 0.6B • Updated May 27 • 84 • 13
Grouped Query Experts: Mixture-of-Experts on GQA Self-Attention Paper • 2606.20945 • Published Jun 18 • 81
FrontiersMind/Nandi-Mini-V1.1-600M-Intermediate-Checkpoint-400GT Text Generation • 0.6B • Updated May 30 • 25 • 9
view article Article How I contributed a new model to the Transformers library using Codex nielsr • Mar 30 • 53
Running on CPU Upgrade 280 The Synthetic Data Playbook: Generating Trillions of the Finest Tokens 📝 280 Visualize synthetic‑data experiments as an interactive bookshelf
view article Article The 1 Billion Token Challenge: Finding the Perfect Pre-training Mix codelion • Nov 3, 2025 • 66