File size: 3,111 Bytes
d4c2896
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
# Final analysis - `experiments/analysis/`

Generated by `scripts/analyze_final.py` (Python 3.12, stdlib only + numpy + matplotlib).

## Methodology

### Sources

- `experiments/runs/<run>/results.jsonl` - one OK record per line (schema in the
  inventory). `config.json` normalizes the run: family (`final/curves/ksweep`),
  target (`{qwen|gemma}-{q4|q5|q8}` from the model path), drafter from
  `spec_type` + `p_min` (dspark) + draft quantization (dflash-f16/q4/q8).
- `server.log` from `curves-*`/`ksweep-*` runs - `draft acceptance` and
  `acc per pos` lines (only place with per-position curves; the `final-*` runs
  rotated logs and their original segment has no acceptance lines).

### Exclusions (timing sentinels)

llama-server reports `tok_per_s = 1,000,000` / `predicted_ms = 0` in 271 records
from Gemma runs (timing quirk, not real speedup). **Rule applied in every
calculation**: exclude any record with `tok_per_s >= 1e5` or `predicted_ms <= 0`
(together with `alpha/tau/draft_n = None`). Qwen has no sentinels. The Gemma
`metrics.json` files are contaminated by this (mean 12,291 tok/s) and are **not**
used for tok/s; only `duration_s` and `errors` are used. Counts per run in
`summary.json` -> `meta.exclusions_per_run` (q4: 18/run, q5: 25/run, q8: 20/run
in final; 3-4/run in gemma curves/ksweep).

### Log <-> record mapping (curves/ksweep)

The order of the `draft acceptance` lines in `server.log` == order of records
with non-None `alpha` in `results.jsonl` (verified in 41/41 runs, max |diff| =
5e-5, rounding only). Gemma sentinel records (alpha=None) have no line and are
skipped. This assigns the domain to each line. No mismatches in any run.

### Speedup vs solo

Per-prompt ratio `tps_draft / tps_solo` matched by `id` against the `solo` run
of the same family+quant; mean and median of the ratio are reported, plus the
ratio of means (`agg`). The 6 solo runs (qwen q4/q5/q8, gemma q4/q5/q8) exist
and are the baseline.

### Break-even alpha_be (paper #32, Bielik et al.)

Per `ksweep-*` run and domain (and aggregated): OLS of `TPS = a + beta*alpha`
over the run records (alpha and per-prompt tok/s, clean records with alpha).
`alpha_be = (TPS_base - a) / beta` where `TPS_base` = mean tok/s of the
corresponding `solo` run (same family+quant+domain, final run). CI95 by delta
method over the OLS covariance of (a, beta). alpha_be > 1 = no reachable
break-even; < 0 = always above baseline. Paper #32 reports k=2: 38.0-52.8% and
k=4: 77.7-90.1% (Table 5); the compact 40-77% range covers k=2..4.

## Files

- `summary.json` - consolidated numbers (per config x domain and breakeven block).
- `tables/t2_speedup.md` ... `tables/t6_breakeven.md` - markdown tables for the paper.
- `curves/acc_by_pos_<run>.csv` - per-position acceptance per domain (41 runs).
- `checks.md` - result of the checks vs session 08.
- `manuscript/figures/F1..F4_*.png` - figures (300 dpi; local paper output, `manuscript/` is not versioned).

## Execution

```
.venv/bin/python scripts/analyze_final.py \
  --runs-dir experiments/runs --out-dir experiments/analysis \
  --figs-dir manuscript/figures
```