glorified-spellcheck / docs /archive /performance-path-forward.md
Chris Cameron
docs: move investigation docs to docs/archive/
55de792
|
Raw
History Blame Contribute Delete
6.89 kB

A newer version of the Gradio SDK is available: 6.22.0

Upgrade

Historical document. Retained as reference for the strategic reassessment after the torch.compile failure and L4 profiling. See docs/lessons-learned.md for the consolidated summary.

Reassessment: best path forward for workshop throughput

Synthesizes the torch.compile failure, the L4 profiling numbers, and the earlier pre-quantized investigation to pick a pragmatic next path.

Updates since written (2026-07-06):

  • ShieldGemma stays eagerly loaded. Lazy-loading was rejected because the safety guard will be used during the workshop; loading everything before the workshop starts is preferable to a mid-workshop cold-start. Mitigation is operational (warm the app once before attendees arrive), not a code change.
  • py-spy is unavailable on Hugging Face Spaces. The GIL-vs-CPU-saturated gate this doc proposes cannot be run there, and is no longer a blocker: horizontal scaling and vLLM both help regardless of the answer, so the distinction is informational only.

What we now know

  • Baseline (compile off): 0.18 req/s under the 15-user/180 s stress test.
  • torch.compile enabled: 0.10 req/s — a ~44 % regression.
  • Cudagraphs are not usable because the app calls compiled models from multiple Gradio worker threads. They only "work" for the fixed-shape next-token forward, and even there they crash.
  • GPU utilization is low (~17 % on T4/L4). The bottleneck remains host-side Python/GIL overhead per token, not GPU kernel throughput or model weight size.
  • Pre-quantized checkpoints exist for gemma-3-1b-it but do not address the diagnosed bottleneck, and they do not solve for base/270m variants.

Conclusion: the root-cause diagnosis in docs/profiling-analysis.md is still correct, but torch.compile is not the fix. We need either (a) remove the GIL contention by isolating each generation, or (b) batch requests so fewer, fatter GPU calls amortize Python overhead.

Decision matrix

Path Expected throughput gain Engineering cost T4 VRAM feasibility Recommendation
Tune config (timeout/concurrency/caps) + lazy ShieldGemma Low (restores reliability) Very low Yes Do this immediately
Process isolation / multiprocessing Medium-High (removes GIL) High Tight with 1B base + instruct + ShieldGemma Fallback if vLLM can't be used
In-process batching High (fewer, fatter calls) Very high (new scheduler + streaming) Yes Highest ceiling but complex
vLLM with continuous batching High (optimized engine) High (replace backend) Yes for 1B models Best long-term path

Recommended path

Immediate (this week): stabilization without compile

Do the minimum to make the current architecture reliable on the L4 (and still deployable on the T4 small):

  1. Set TORCH_COMPILE=False (or disable the feature flag) in production.
  2. Lazy-load ShieldGemma so the ~5.2 GB safety model is only loaded when the safety guard is toggled on. This was excluded earlier but is now high value and low risk.
  3. Tune Gradio concurrency on the L4/T4:
    • Start with GRADIO_CONCURRENCY=2.
    • Stress-test at 2, 3, 4, and 5 concurrent requests under the realistic workload.
    • Pick the value that maximizes successful requests per second.
  4. Cap response length for the demo:
    • Lower DEFAULT_MAX_TOKENS and DEFAULT_BASE_MAX_TOKENS to 128 or 100.
    • This shortens every generation and therefore raises throughput.
  5. Optional: use MODEL_SIZE=270m if the demo content allows it. It is already supported and drastically cheaper.

Target outcome: a stable, predictable workshop at the cost of some output length and parallel capacity.

Short term (next iteration): prototype a vLLM backend

For a real throughput improvement, the highest-leverage change is replacing the in-process Transformers backend with vLLM:

  • vLLM uses continuous batching and paged attention, which directly solves the "many tiny GPU calls" problem.
  • It handles streaming generation and batched generation out of the box.
  • Two vllm.LLM engines (one base, one instruct) fit comfortably on a T4 or L4 for 1B/270M models.
  • It does not require torch.compile to be fast.

This is a contained backend replacement: llm_backend/generation.py, llm_backend/logit_pipeline.py, and llm_backend/safety.py. The Gradio UI in app.py stays mostly the same because the public function signatures (generate_comparison_response, generate_single_response, predict_next_token) can be preserved.

Contingent: process isolation

If vLLM is not an option (e.g., deployment restrictions, dependency issues on Hugging Face Spaces, or the team does not want to rewrite the backend), then process isolation is the next best alternative:

  • Run each model variant in its own process.
  • Each process has its own Python interpreter and GIL, eliminating GIL contention.
  • This is the path already described in the main optimization plan’s Task 3.
  • On a T4, load only what fits; on the L4, load base and instruct in two long-lived processes and keep ShieldGemma lazy.

Why not process isolation first?

Process isolation would remove the GIL contention, but it:

  • Adds significant IPC and streaming complexity for a modest model size.
  • Requires either duplicating model weights in memory (one process per variant) or implementing shared-memory sharing.
  • Does not solve the "too many tiny calls" problem; it only makes those calls contend less.

vLLM solves both the GIL/contention problem and the tiny-kernel problem, so it is the better investment of engineering time.

Pre-quantized checkpoint role

Pre-quantized gemma-3-1b-it checkpoints (e.g., RedHatAI/gemma-3-1b-it-quantized.w8a8) become relevant again only if we move to vLLM and need to squeeze memory on the T4 small. They are not useful for the current Transformers-only backend. Defer decision until the vLLM prototype is running.

Proposed next steps (ordered)

  1. Merge or apply the stabilization changes (disable compile, lazy ShieldGemma, concurrency tuning).
  2. Run the 15-user stress test on the L4 to get a new baseline.
  3. Start a branch to spike a vLLM backend:
    • Load google/gemma-3-1b-pt and google/gemma-3-1b-it as two vLLM engines.
    • Implement generate_single_response and generate_comparison_response using vLLM synchronous/async streaming APIs.
    • Keep predict_next_token and check_safety functional.
  4. Stress-test the vLLM branch under the same 15-user load.
  5. Compare vLLM req/s, success rate, and VRAM against the stabilized baseline.

If vLLM delivers a clear win, adopt it and drop the process-isolation plan. If vLLM is blocked by deployment issues, fall back to the process-isolation plan.