Serving Economics
Inference cost decides your architecture. These papers, read top to bottom, build the mental model: where the FLOPs and the bytes actually go.
Paper • 2211.05102 • Published • 4Note Start here. Pope et al. give you the actual arithmetic — memory bandwidth vs. FLOPs, why decode is bandwidth-bound and prefill is compute-bound. Every decision below follows from this.
Scaling Laws for Neural Language Models
Paper • 2001.08361 • Published • 11Note Scaling laws. The baseline everyone quotes; read it to know what it does *not* say about inference.
Training Compute-Optimal Large Language Models
Paper • 2203.15556 • Published • 13Note Chinchilla. Compute-optimal training gives you smaller models for the same quality — which is really an inference-cost result.
Fast Transformer Decoding: One Write-Head is All You Need
Paper • 1911.02150 • Published • 9Note Multi-query attention. Shazeer's short paper is the origin of every KV-cache reduction that followed.
GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints
Paper • 2305.13245 • Published • 6Note GQA. The interpolation between MQA and MHA that essentially every modern open model now ships.
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Paper • 2205.14135 • Published • 16Note FlashAttention. IO-awareness as a design principle — the same reasoning applies far beyond attention.
FlashAttention-2: Faster Attention with Better Parallelism and Work Partitioning
Paper • 2307.08691 • Published • 10Note FlashAttention-2. Read for the work-partitioning discussion; it is a lesson in occupancy.
Efficient Memory Management for Large Language Model Serving with PagedAttention
Paper • 2309.06180 • Published • 66Note PagedAttention / vLLM. Virtual memory applied to the KV cache. The single highest-leverage serving idea of the last few years.
Fast Inference from Transformers via Speculative Decoding
Paper • 2211.17192 • Published • 11Note Speculative decoding. Trades cheap parallel compute for expensive sequential memory traffic. Know the acceptance-rate math before you promise a speedup.
Efficiently Programming Large Language Models using SGLang
Paper • 2312.07104 • Published • 8Note SGLang / RadixAttention. Prefix sharing across requests — huge for agent and few-shot workloads with common prefixes.
Mooncake: A KVCache-centric Disaggregated Architecture for LLM Serving
Paper • 2407.00079 • Published • 6Note Mooncake. Disaggregated prefill/decode in production. This is the shape serving fleets are converging on.
H_2O: Heavy-Hitter Oracle for Efficient Generative Inference of Large Language Models
Paper • 2306.14048 • Published • 14Note H2O. KV-cache eviction — the beginning of treating cache entries as unequal.
Efficient Streaming Language Models with Attention Sinks
Paper • 2309.17453 • Published • 15Note StreamingLLM / attention sinks. Why naive KV eviction collapses, and the cheap fix.
Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity
Paper • 2101.03961 • Published • 13Note Switch Transformers. MoE from the systems side: capacity factors and routing cost, not just the quality win.
Mixtral of Experts
Paper • 2401.04088 • Published • 162Note Mixtral. The report that made sparse MoE the default assumption for open frontier models.
DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model
Paper • 2405.04434 • Published • 29Note DeepSeek-V2. Multi-head latent attention — an aggressive KV-cache compression that actually shipped at scale.
axjns/strix-halo-inference-bench
Viewer • Updated • 44 • 27 • 1Note The papers above, reduced to numbers on one fully-specified machine. Prefill vs decode on a bandwidth-limited unified-memory part, measured rather than argued.