LLM Compression by Block Removal with Constrained Binary Optimization Paper • 2602.00161 • Published Jun 17 • 9
view article Article Pruning LLMs Like a Physicist: Block Removal as an Ising Optimization Problem MultiverseComputingCAI • 5 days ago • 31
Compile by Training: Turning Natural-Language Specifications into Local Neural Functions Paper • 2609.04199 • Published 24 days ago • 328
LLaDA-Image: Building Strong Image Generators with Fully Open Training Recipes Paper • 2609.03796 • Published 24 days ago • 186
view article Article Fine-tuning a 350M Model for Better Structured Outputs in 100 GRPO Steps +1 iamleonie, burtenshaw, sergiopaniego • 24 days ago • 137
view article Article LFM2.5 Q4\_0 Checkpoints from Quantization-Aware Distillation LiquidAI • Aug 19 • 36
view reply Have you experimented with or tested different vocabulary sizes to see if the vocab size directly impacts the overtraining threshold and benchmark peak?