Running 62 Don't Train the Model, Evolve the Harness 🌿 62 Evolving an agent's harness, not its model, on Harvey's LAB
Running 223 The ultimate guide to RL environments: building and scaling them in the LLM era 📝 223 Building and scaling RL environments for LLM training
Running Featured 92 Distilling 100B+ Models 40x Faster with TRL 📝 92 TRL distillation for 100B+ teachers, 40x faster
Running on CPU Upgrade 270 The Synthetic Data Playbook: Generating Trillions of the Finest Tokens 📝 270 Visualize synthetic‑data experiments as an interactive bookshelf
Running Featured 83 QED-Nano: Teaching a Tiny Model to Prove Hard Theorems 📝 83 Who needs 1T parameters? Olympiad proofs with a 4B model