glorified-spellcheck / docs /lessons-learned.md
Chris Cameron
docs(lessons-learned): record ZeroGPU workshop load-capacity results
c35d5dc
|
Raw
History Blame Contribute Delete
8.8 kB
# Lessons Learned — Dead Ends & Disproven Approaches
> **Read this before pursuing performance or architecture changes.**
> Each entry records what was tried, why it failed, and when it might be
> worth revisiting. Detailed reasoning lives in `docs/archive/`.
---
## Transformers + dedicated GPU for workshop load
**Status:** Disproven for >1 concurrent user.
**Outcome:** 0.18 req/s, 25% success rate at 15 concurrent users on a dedicated T4. GPU utilization stuck at ~17%.
**Why it failed:**
- Python/GIL contention in the `model.generate()` loop serializes threads under concurrent load.
- Per-token Python overhead (logit processing, sampling, prompt re-packing) dominates; CUDA kernels are fast but the host can't feed them fast enough.
- Each additional Gradio worker thread makes every other thread slower, until long generations stall past the timeout.
**Detail:** `docs/archive/profiling-analysis.md` §3-5.
**When to revisit:** Never, unless the generation loop moves out of Python (vLLM, process isolation, or a compiled backend that actually works under multithreading).
---
## torch.compile
**Status:** Disproven — 44% throughput regression.
**Outcome:** 0.10 req/s with compile on vs 0.18 req/s off (L4, 15-user stress test).
**Why it failed:**
- Cudagraphs crash under Gradio's multithreaded handlers — Inductor stores cudagraph tree state in thread-local storage, so worker threads trip an `AssertionError`. Without cudagraphs, `reduce-overhead` loses most of its benefit for tiny per-token kernels.
- `dynamic=True` + varying prompt lengths means each new shape pays recompilation cost; at classroom scale those first calls dominate.
- Models are so small (270m-1B) that kernel execution time is already tiny; Dynamo guard + Inductor codegen dispatch overhead outweighs fusion benefit.
**Detail:** `docs/archive/profiling-analysis.md` §"torch.compile measured regression" addendum.
**When to revisit:** Only if the threading model changes (single-threaded inference) or the model size grows large enough that kernel fusion matters.
---
## Quantization (INT8 / AWQ / FP8)
**Status:** Disproven — wrong bottleneck.
**Outcome:** GPU is already idle at 17% utilization. Reducing weight size/compute does not help when the bottleneck is host-side Python overhead.
**Why it failed:**
- Quantization reduces memory bandwidth and VRAM footprint, but the GPU is not memory-bound or compute-bound — it's starved by the Python dispatch loop.
- Pre-quantized checkpoints exist only for some model variants (1B instruct); no coverage for base or 270m.
- Community checkpoints introduce quality, tokenizer, and dependency risks.
**Detail:** `docs/archive/pre-quantized-models-assessment.md`.
**When to revisit:** Only as part of a vLLM migration where quantized weights are paired with a continuous-batching engine that can exploit them.
---
## Raising GRADIO_CONCURRENCY on Transformers
**Status:** Disproven — triggers GIL collapse.
**Outcome:** At concurrency=4, 75% failure rate (all long-generation requests timed out). At concurrency=1, 100% success but serial.
**Why it failed:**
- GIL contention is not a mild slowdown — it's a catastrophic collapse. Each additional thread makes every thread slower until long generations stall past the 30s timeout.
- The correct fix is to reduce per-request server time, not increase concurrency.
**Detail:** `docs/archive/profiling-analysis.md` §5.
**When to revisit:** Never on the Transformers backend. vLLM's continuous batching handles concurrency without Python-level contention.
---
## CUDA_LAUNCH_BLOCKING=1
**Status:** Disproven — made things 10x worse.
**Outcome:** Next-token forward pass went from 0.18s to ~2.0s. All streaming requests timed out. GPU util did not rise.
**Why it failed:**
- Blocking removes the CPU/GPU overlap that was partly masking the GIL problem.
- A 270m forward/decode is hundreds of tiny kernels; blocking forces a CPU↔GPU sync on each one.
- GPU util stayed flat at ~10-17%, confirming the GPU was not secretly busy.
**Detail:** `docs/archive/profiling-analysis.md` §4.
**When to revisit:** Never. The diagnosis is closed: async dispatch was helping, not hurting.
---
## vLLM on T4 (Turing)
**Status:** Incompatible.
**Outcome:** vLLM 0.24 V1 engine requires Ampere+ (compute 8.0). T4 is Turing (compute 7.5).
**Why it failed:**
- V1 engine requires Ampere+ for bf16, FlashAttention 2, and Triton shared memory.
- No workaround; T4 is fundamentally below the floor.
**Detail:** `docs/archive/vllm-spike-results.md` §"Hardware note".
**When to revisit:** Only if vLLM adds Turing support (unlikely) or you move to Ampere+ hardware.
---
## vLLM on CPU (0.24 pinned)
**Status:** Not viable with current pin.
**Outcome:** The standard `vllm==0.24.0` wheel is GPU-only.
**Why it failed:**
- vLLM's CPU backend is experimental with no official prebuilt wheels.
- Requires a third-party `vllm-cpu` PyPI package or a source build.
- Needs AVX512-class instructions (`avx512_bf16` / `avx512_vnni`) for usable performance; plain x86-64 crawls.
- The `AsyncLLMEngine` API used in this repo is untested on CPU.
**Detail:** This spec's research (2026-07-08).
**When to revisit:** If vLLM CPU support matures to official wheels and the `AsyncLLMEngine` path is validated on CPU.
---
## ZeroGPU + vLLM
**Status:** Fundamentally incompatible.
**Why it fails:**
- vLLM engines are long-lived processes that grab VRAM at construction and hold persistent state.
- ZeroGPU allocates GPU per-call from a shared pool with short-lived allocation windows.
- They cannot coexist architecturally.
**Detail:** `AGENTS.md` history; commit `55d2416`.
**When to revisit:** Never. This is an architectural incompatibility, not a version issue.
---
## Process isolation / multiprocessing
**Status:** Considered, deprioritized.
**Outcome:** Would remove GIL contention but at high complexity for modest gain.
**Why it was passed over:**
- Adds IPC and streaming complexity for small models.
- Requires duplicating model weights per process or implementing shared-memory sharing.
- Does not solve the "too many tiny kernel calls" problem — only makes them contend less.
- vLLM solves both the GIL/contention problem AND the tiny-kernel problem, making it the better investment.
**Detail:** `docs/archive/performance-path-forward.md`.
**When to revisit:** If vLLM is blocked by deployment restrictions and throughput must improve.
---
## ZeroGPU workshop load capacity
**Status:** Resolved — measured 2026-07-12. ZeroGPU is viable for a ~15-user workshop.
**Run:** `scripts/stress_test.py --users 15 --mode realistic --duration 180` against the deployed Space (`rectified-snugness/glorified-spellcheck`, ZeroGPU H200, Transformers backend). 100 requests / 187 s wall.
**Outcome:** **0.53 req/s, 100% success** (0 timeouts, 0 errors, 0 rate-limited). Per-endpoint client latency:
| Endpoint | p50 | p95 | p99 |
|---|---|---|---|
| `next_token` | 5.4 s | 12.8 s | 14.8 s |
| `compare_models` | 8.6 s | 28.6 s | 32.7 s |
| `layering` | 8.6 s | 33.0 s | 37.9 s |
**Interpretation:**
- **Reliability fixed vs the T4 baseline** (`docs/archive/profiling-analysis.md`): 25% → 100% success and 0.18 → 0.53 req/s (~3×). The H200 is fast enough that even GIL-contended generations finish within `DEFAULT_TIMEOUT`, so the collapse that crippled the T4 no longer causes failures.
- **The GIL still caps throughput** at ~0.53 req/s — roughly 7× below vLLM's dedicated-L4 ceiling (3.63 req/s, `docs/archive/vllm-spike-results.md`). ZeroGPU did not remove the Python bottleneck; it stopped it from breaking requests. vLLM remains a ZeroGPU dead-end (dedicated GPU only), so this gap is the accepted trade-off.
- **`next_token` latency is wait-bound**, not compute-bound: its forward pass is ~21 ms warm but p50 is 5.4 s — the gap is queueing behind slow generator calls plus ZeroGPU's per-call GPU-grant overhead. On a dedicated GPU this lightweight endpoint interleaved cheaply (0.5 s p50); on ZeroGPU every `@GPU` call pays a fixed allocation cost that dominates a 21 ms forward pass.
**Reproduce:** `scripts/stress_test.py --url <owner/name-or-*.hf.space-app-url> --users 15 --mode realistic --duration 180` (not the `huggingface.co/spaces/...` Hub page URL — `gradio_client` cannot fetch config from it).
**When to revisit:** Before any workshop larger than ~15 users, or if p95 latency (~30 s on the generators) becomes painful. This single run exceeds the free tier's 300 GPU-s/day budget; a PRO account (2400 s/day) absorbs it with headroom for a few repeats. If throughput must improve, the only remaining lever is vLLM on a dedicated GPU (`archive/vllm-backend` branch + `docs/archive/vllm-*.md`); the GIL ceiling cannot be raised on ZeroGPU.