Running 4.08k The Ultra-Scale Playbook 🌌 4.08k The ultimate guide to training LLM on large GPU Clusters
Running on CPU Upgrade Featured 3.32k The Smol Training Playbook 📚 3.32k The secrets to building world-class LLMs
AdaPerceiver: Transformers with Adaptive Width, Depth, and Tokens Paper • 2511.18105 • Published Nov 22, 2025
AdaPerceiver-V1 Collection This is the collection of AdaPerceiver models. • 2 items • Updated Dec 22, 2025