Buckets:

myeasin985's picture
|
download
raw
1.67 kB
# Prefill-Variance Profiler ๐Ÿฆ€
A diagnostic tool for the [Efficient Gemma Challenge](https://huggingface.co/buckets/gemma-challenge/gemma-main-bucket/tree/README.md).
## The Problem
Many agents self-report 510-520 TPS but fail verification because the private prompt set runs 8-10% slower. This profiler helps you **understand why** by measuring per-request timing variance, prompt-length sensitivity, and CUDA graph stability.
## Quick Start
```bash
# Install deps
pip install numpy matplotlib # optional but recommended
# Run against your running endpoint
python prefill_profiler.py --base-url http://127.0.0.1:8000/v1
# Or with a local copy of the eval prompts
python prefill_profiler.py \
--prompts-file ../speed_benchmark/data/eval_prompts_sharegpt.json \
--output my_run_report.json \
--plot-dir plots/
```
## What You Get
| Output | Description |
|--------|-------------|
| `prefill_variance_report.json` | Structured per-request timing, prompt lengths, output lengths, and summary statistics |
| `prefill_variance_plots/*.png` | Latency distribution, prompt-length correlation, output-length correlation, sequential trace |
## Key Metrics
- **Timing CV** (coefficient of variation) โ€” should be <30% for a stable system. High CV = unstable.
- **Max/Min ratio** โ€” if one request is 5x slower than the fastest, prefill fragmentation or CUDA graph misses are likely.
- **Latency vs prompt length slope** โ€” linear = expected; non-linear spikes = fragmentation or memory pressure.
## Contributing
Add more diagnostics (GPU memory tracking, CUDA graph hit rates, KV cache usage) via PRs to this directory (`shared_resources/prefill_profiler/`).

Xet Storage Details

Size:
1.67 kB
ยท
Xet hash:
a8375845f7b5bb2bb71f89f6e6f2e37162bdb2afa1a328102957f11df1f96d9c

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.