Buckets:
Benchmark Observations & Notes
Chef et Dev | Gemma Speed Challenge
Infrastructure Notes
Hardware: NVIDIA A10G 24GB
- Compute capability: 8.9
- Memory bandwidth: ~378 GB/s
- FP16/BF16 performance: ~156 TFLOPS
- FP8 performance: ~313 TFLOPS (sparse)
Known Issues Encountered
- Volume mount failures on HF Jobs (resolved)
- rope_scaling compatibility with vLLM 0.7.3 (resolved by upgrading to 0.22.0)
--no-enable-log-requestsflag removed in newer vLLM (resolved)- FlashInfer not installed in harness image (resolved)
vLLM 0.22.0 Default Configuration
# Default from harness serve.py:
--model google/gemma-4-E4B-it
--served-model-name gemma-4-e4b-it
--host 0.0.0.0
--port 8000
--dtype bfloat16
--max-model-len 4096
--gpu-memory-utilization 0.90
--trust-remote-code
--no-enable-log-requests
--max-num-batched-tokens 512 # (if provided)
What This Means
- BF16 inference — standard precision, no quantization yet
- 4096 context — focused on single-stream, reasonable context length
- 90% GPU memory — leaves 10% for KV cache + PPL stage
- 512 batched tokens — bounds prefill chunk size for PPL stability
Expected Performance Tiers
| Tier | TPS | Technique |
|---|---|---|
| Tier 1 | ~300-320 | vLLM default BF16 (vanilla) |
| Tier 2 | ~350-380 | vLLM + FlashInfer + CUDA graphs |
| Tier 3 | ~380-400 | FP8 quantization + optimized vLLM |
| Tier 4 | ~400-416 | AWQ/FP8 + kernel optimizations |
| Frontier | ~416+ | Aggressive combined optimizations |
Optimization Research Priorities
High Impact (Quick Wins)
- FlashInfer installation — should be automatic with vLLM 0.22.0
- CUDA graph warmup — ensure stable capture before timing
- PagedAttention tuning — max_seq_len_to_capture parameter
Medium Impact
- FP8 quantization — native A10G support
- KV cache quantization — FP8 KV cache
- Scheduler tuning — delay_factor, num_scheduler_steps
Lower Impact
- AWQ int4 — requires weight conversion
- Triton kernels — manual implementation needed
- FlashAttention-3 — needs custom install
PPL Stage Requirements (MUST comply)
- Endpoint must support
/v1/completions - Must accept
promptas list of int token IDs - Must accept
prompt_logprobs - Must respect
add_special_tokens: false - Must return
prompt_logprobsin response - Must support
return_token_ids: true - Greedy decode must match reference exactly
Last updated: 2026-06-11 by chef-et-dev
Xet Storage Details
- Size:
- 2.56 kB
- Xet hash:
- 6f813505d01c8ca1ef21baf2cdfec99d8275402b9597711e5e0525fdd895108d
·
Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.