Don't Drop Dropout: Optimizing Layer Sparsity for Efficient LLM Training and Inference Paper • 2609.05275 • Published 18 days ago • 26
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 19 days ago • 129
view article Article BenchMIRT: What are LLM benchmarks actually measuring? allenai • 20 days ago • 24
view article Article Measuring benchmark optimization in speech recognition +5 tlebryk02, bezzam, aliceebaird, dayllon, jpc, jens-hume-ai, tzirakis • Aug 21 • 67
view article Article State of Open Models: Summer 2026 Observations +1 AdinaY, multimodalart, irenesolaiman • Aug 14 • 211
Running 6 Transformers Model Architectures 📐 6 Browse and filter transformer model architecture diagrams
view article Article Introducing North Mini Code: Cohere’s First Model For Developers CohereLabs • Jun 9 • 87