The Expert-Cache Cliff
📉
Why LRU hits exactly zero in mixture-of-experts inference
Quantization as cache amplification: expert offloading, sub-2-bit codecs and cache policy for MoE inference on a single laptop.
Why LRU hits exactly zero in mixture-of-experts inference
Note Live in-browser simulator: watch an LRU expert cache return exactly 0% below the per-token working set.
Note The paper, full source, raw routing traces and every measurement artefact behind every number.
Note The MoE studied throughout: 6.9B total / 1.3B active, 16 layers, 64 experts/layer, top-8.