# Piko-9b — Performance Report Every number here was measured on the hardware listed below. Nothing is extrapolated to other GPUs, other precisions, or other context lengths. Configurations that failed are recorded as failures, not omitted. Raw data: [`benchmarks/results/inference_4bit.json`](../benchmarks/results/inference_4bit.json), [`benchmarks/results/memory_4bit.json`](../benchmarks/results/memory_4bit.json). ## Test environment | | | |---|---| | GPU | NVIDIA GeForce RTX 5070 Ti, 15.92 GB usable | | CPU RAM | 33.9 GB | | OS | Windows 11 + WSL2 (Linux 6.6.87.2-microsoft-standard) | | Python | 3.12.3 | | torch | 2.10.0+cu128 | | transformers | 5.5.0 | | Precision | bfloat16 compute, **4-bit NF4** weights (bitsandbytes, double quant) | | Placement | `device_map={"": 0}` — fully resident, no CPU offload | | Decoding | greedy (`do_sample=False`) | | Linear-attention kernels | **not installed** — pure PyTorch fallback | | Weights location | internal NVMe | Because `flash-linear-attention` and `causal-conv1d` were not installed, every throughput figure below is a **floor**. Installing them should improve it; by how much was not measured. ## Cold load | Measurement | Value | |---|---| | Load time, NVMe | **101–119 s** | | Load time, external USB via WSL 9P | **10–25 min** (observed repeatedly) | | Resident VRAM after load | **7.37 GB** | | Host RSS before load | 0.55–0.81 GB | | Host RSS after load | **1.28 GB** | The 10× difference between NVMe and USB dominates every other cost in this report. Copy the checkpoint to internal storage before doing repeated work. `reported_parameters` from the loaded model is 5,724,972,272 — this is the **packed** element count under 4-bit quantization, not the logical parameter count. The true figure, read from the safetensors headers, is **9,653,104,368**. ## Latency and throughput `max_new_tokens=128`, greedy. TTFT is a full `generate()` capped at one new token; the decode rate excludes that first token. Both are measured after a warm-up pass. | Context | Batch | TTFT (s) | Prefill (tok/s) | Decode (tok/s) | Peak VRAM (GB) | |---:|---:|---:|---:|---:|---:| | 512 | 1 | 0.139 | 3,689 | **29.1** | 7.53 | | 512 | 2 | 0.230 | 4,461 | **42.8** | 7.62 | | 512 | 4 | 0.411 | 4,986 | **84.9** | 7.83 | | 2,048 | 1 | 0.407 | 5,036 | **35.8** | 7.75 | | 2,048 | 2 | 0.750 | 5,466 | **43.2** | 8.15 | | 2,048 | 4 | 1.470 | 5,576 | **82.6** | 8.95 | | 8,192 | 1 | 1.475 | 5,553 | **34.8** | 8.86 | | 8,192 | 2 | 2.925 | 5,601 | **42.4** | 10.38 | | 8,192 | 4 | 5.866 | 5,584 | **81.2** | 13.40 | All nine configurations completed. No failures. Three things worth noting: **Decode rate is nearly flat across context length.** 29–36 tok/s at batch 1 whether the prompt is 512 or 8,192 tokens. That is the hybrid architecture working as designed: 24 of 32 layers keep a fixed-size recurrent state, so only the 8 full-attention layers pay a growing KV cost. **Batching scales well.** Batch 4 gives ~2.4–2.9× the aggregate throughput of batch 1 at a modest memory cost. **Prefill throughput is stable at ~5,000–5,600 tok/s** once past the shortest prompts. ## Memory by context length Single forward pass with `use_cache=True`. "Added" is peak minus the resident weights. | Context | Peak VRAM (GB) | Added over weights (GB) | Prefill (s) | |---:|---:|---:|---:| | 512 | 7.63 | +0.29 | 0.805 | | 2,048 | 8.40 | +1.05 | 0.402 | | 8,192 | 11.48 | +4.13 | 1.590 | | 32,768 | **failed** | — | — | **32,768 tokens failed** on this GPU with `RuntimeError: CUDA driver error: device not ready` — an out-of-memory condition surfacing as a driver fault. Extrapolating the 8K figure (+4.13 GB) puts a 32K single forward pass well beyond the 15.92 GB available. This is a limit of **this GPU**, not of the model. It is also narrower than what generation can do: the custom suite's long-context cases at ~32,000 filler tokens **passed**, because `generate()` builds its cache incrementally rather than materialising a full-length forward pass at once. So, precisely: * 32K-token retrieval through `generate()`: **works** on 15.92 GB. * 32K-token single forward with full cache retention: **does not fit** on 15.92 GB. ## Image preprocessing | Measurement | Value | |---|---| | Median | **13 ms** | | Min / max over 5 repeats | 10.1 ms / 22.3 ms | | Prompt tokens for a 520×300 receipt | 214 | Negligible against generation time. A 520×300 document consumes ~200 tokens of context. Measured end-to-end image + text answers during validation: **~4 s**, against 0.9–2.9 s for a comparable text-only answer. ## Real-world latencies from the validation run Taken from [`reports/inference_validation.json`](inference_validation.json), 4-bit, NVMe: | Task | Time | |---|---| | Short text answer (48 tokens) | 0.85–1.5 s | | Code generation (96 tokens) | 2.9 s | | Image OCR / document JSON (128 tokens) | ~4.0–4.3 s | | Multi-turn follow-up (32 tokens) | 0.91 s | | Batch of 2, short answers | 1.19 s | | 14,429-token prompt + 32-token answer | 3.53 s | ## Configurations not measured | Configuration | Status | Why | |---|---|---| | bfloat16 unquantized | **Not run** | Needs ~22 GB; this GPU has 15.92 GB | | 8-bit | **Not run** | Would fit, but was not exercised | | Context 32K single forward | **Failed** | OOM on this GPU; recorded above | | Context beyond 32K | **Not run** | Hardware limit | | Multi-GPU | **Not run** | Splitting risks the same corruption as CPU offload | | CPU-only | **Not run** | ~38.6 GB float32 and no fast path; not viable | | With linear-attention kernels | **Not run** | All figures are therefore a floor | | vLLM / SGLang serving | **Not run** | Engine support for `qwen3_5` unverified | ## Reproducing ```bash python benchmarks/profile_inference.py \ --model --quantization 4bit \ --context-lengths 512 2048 8192 --batch-sizes 1 2 4 --max-new-tokens 128 \ --image evaluation/custom_suite/assets/receipt.png \ --output benchmarks/results/inference_4bit.json python benchmarks/profile_memory.py \ --model --quantization 4bit \ --context-lengths 512 2048 8192 32768 \ --output benchmarks/results/memory_4bit.json ``` Roughly 20 minutes total on comparable hardware, plus load time. ## A methodology note An earlier run of `profile_inference.py` reported decode rates of 33 and 164 billion tokens per second. The cause was measuring prefill as a bare forward pass and subtracting it from the full `generate()` time: when the subtraction went negative it was clamped to a floor, and the division exploded. The script now measures TTFT as `generate(max_new_tokens=1)`, warms up first, and reports `null` with a note when the decode window is too short to measure rather than emitting a number. The bad figures were never published; they are recorded here because a benchmark harness that can produce absurd values silently is worth documenting.