File size: 11,375 Bytes
a4508e8
 
 
 
3c35b12
 
 
 
a4508e8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3c35b12
 
a4508e8
 
 
3c35b12
 
 
 
 
 
 
a4508e8
3c35b12
 
a4508e8
 
 
 
3c35b12
 
 
 
 
 
 
 
 
 
a4508e8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3c35b12
 
 
 
 
a4508e8
 
 
 
 
 
 
 
3c35b12
 
 
df300c9
3c35b12
 
 
a4508e8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
3c35b12
a4508e8
 
 
 
 
 
df300c9
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
# GPU Perf Prophet — Data Card v0.1

> **Status:** Hand-authored baseline. This file will be superseded once `build_data_card()`
> is implemented and `make data` is run.
> §1, §2's MLPerf figures, and §4 are verified against `data/processed/mlperf_raw.parquet`
> on 2026-06-12. §2 also covers the AMD Dev Cloud calibration rows added 2026-07-07, which
> are a separate provenance layered on top of the MLPerf corpus at the training-features
> stage (`data/processed/mlperf_features.parquet`) — they never enter `mlperf_raw.parquet`.

---

## 1 · Provenance

| Field | Value |
|-------|-------|
| Primary source | MLCommons MLPerf Inference results repos v4.1, v5.0, v5.1, v6.0 |
| Divisions | closed + open |
| Ingestor | `src/data/mlperf_parser.py` |
| GPU spec enrichment | `src/data/gpu_spec_db.py``data/gpu_specs.yaml` |
| Processed artifact | `data/processed/mlperf_raw.parquet` |
| Generated | 2026-06-12 |
| Total parsed rows | **1,223** |
| Unique submitters | 32 |
| Unique system configs | 167 |
| Result validity | 100% VALID (zero INVALID rows in corpus) |

Source repos are sparse-checked-out git clones. Fetch scripts in `scripts/fetch_mlperf.sh`.
Every parsed row carries `log_path` back to its originating `mlperf_log_summary.txt`.

---

## 2 · In-scope training corpus

Rows entering model training must pass the `gpu_in_model_scope == True` filter
(set in `data/gpu_specs.yaml`). CPU-only, heterogeneous, edge, and other
out-of-scope GPUs are excluded.

**Total MLPerf in-scope rows: 649** (2026-06-12). Plus 24 self-run AMD Dev Cloud MI300X
calibration rows added 2026-07-07 — **total current in-scope rows: 673.**

### Per-GPU breakdown

| Canonical ID | GPU | Vendor | Rounds present | MLPerf rows | Calibration rows | Total | Spec confidence |
|---|---|---|---|---:|---:|---:|---|
| `mi300x` | AMD Instinct MI300X | AMD | v4.1, v5.0, v5.1, v6.0 | 56 | 24 | **80** | verified |
| `mi325x` | AMD Instinct MI325X | AMD | v5.0, v5.1, v6.0 | 82 | 0 | 82 | verified |
| `mi355x` | AMD Instinct MI355X | AMD | **v6.0 only** | **50** | 0 | **50** | estimated |
| `h100_sxm` | NVIDIA H100 SXM5 | NVIDIA | v4.1, v5.0, v5.1 | 178 | 0 | 178 | verified |
| `h200_sxm` | NVIDIA H200 SXM | NVIDIA | v4.1, v5.0, v5.1 | 283 | 0 | 283 | verified |

AMD total: **212 rows** (31 %)  
NVIDIA total: **461 rows** (69 %)

The AMD/NVIDIA imbalance is the reason the AMD-specific MAPE target is relaxed to
< 20 % (vs. < 15 % overall).

**Calibration provenance:** the 24 MI300X rows were self-run on the AMD Developer Cloud
(vLLM 0.23.0, ROCm 7.2.4) covering `gptj`, `llama2-70b`, `llama3.1-8b`, and `mixtral-8x7b``gptj` and `llama3.1-8b` had **zero** official MLPerf MI300X submissions before this.
These rows measure a genuinely different regime from official submissions (no
serving-stack tuning — see `README.md` § Recommendation accuracy); model-quality metrics
reported anywhere in this project score against official MLPerf rows only, with
calibration rows used purely as additional training signal, never as evaluation ground
truth.

### Per-round in-scope row counts (MLPerf only; excludes calibration rows)

| Round | mi300x | mi325x | mi355x | h100_sxm | h200_sxm | Round total |
|-------|-------:|-------:|-------:|---------:|---------:|------------:|
| v4.1 | 16 | — | — | 116 | 68 | 200 |
| v5.0 | 4 | 16 | — | 54 | 156 | 230 |
| v5.1 | 28 | 62 | — | 8 | 59 | 157 |
| v6.0 | 8 | 4 | **50** | — | — | 62 |
| **Total** | **56** | **82** | **50** | **178** | **283** | **649** |

---

## 3 · MI355X — v6.0-only GPU

**MI355X is the only in-scope GPU with data from a single round.**  
This is the primary known limitation of the training corpus.

### What "v6.0-only" means

| Property | Detail |
|---|---|
| First MLPerf round | v6.0 (no MI355X rows in v4.1, v5.0, or v5.1) |
| In-scope training rows | 50 |
| Benchmarks covered | `llama2-70b-99` (25 rows) and `llama2-70b-99.9` (25 rows) only |
| Benchmarks *not* covered | `gptj`, `mixtral-8x7b`, `llama3.1-405b` |
| Scenarios | Offline and Server |
| Spec confidence | **estimated** — CDNA4 dense FLOPS derived from with-sparsity ÷ 2 |

### Why this matters for the model

1. **No round-over-round variance.** The model cannot check MI355X consistency across
   rounds, making it harder to detect submission anomalies or performance regressions.

2. **Single benchmark family.** The model sees MI355X only on `llama2-70b`. Predictions
   for MI355X on `gptj`, `mixtral-8x7b`, or `llama3.1-405b` are extrapolations,
   not interpolations. SHAP explanations will not decompose MI355X benchmark-level
   effects from GPU-level effects for those benchmarks.

3. **50 rows < 100-row minimum.** If MI355X does not meet the 100-row minimum gate
   after further AMD Dev Cloud calibration, it will be marked `enabled: false` in
   `gpu_specs.yaml` and excluded from v1 recommendations. The current value is
   `in_model_scope: true` as a placeholder pending that gate. **Until that gate is
   built, the API/UI discloses the shortfall per-request instead:** MI300X, MI325X,
   and MI355X (all under the 100-row floor) get `training_data_tier: "below_floor"`
   rather than being presented with the same confidence as a well-covered GPU —
   an interim, per-request disclosure, not a substitute for the hard gate above.

4. **Estimated spec values.** `gpu_specs.yaml` MI355X entries for `peak_tflops` are
   `sparse/2` derivations (flagged `spec_confidence: estimated`). These **must be
   reconfirmed** against the AMD MI350-series whitepaper before the training data
   is finalized.

### Mitigation plan

| Step | Action | Status |
|------|--------|--------|
| 1 | Train with MI355X included; report MI355X MAPE separately; note in model card | Done |
| 2 | Run MI300X + MI355X benchmarks on AMD Dev Cloud; add calibration rows | MI300X run 2026-07-07, but under its own ≥50-row / ≥3-LLM / ≥3-batch-size / 2-precision target: delivered 24 rows, 4 LLMs, 2 precisions, **zero batch-size variation**. Meets 2 of 4 axes. MI355X not yet run — Dev Cloud access covers MI300X instances only |
| 3 | If MI355X row count ≥ 100 → keep `enabled: true`; else set `enabled: false` | Still blocked — MI355X remains at 50 rows |
| 4 | Public writeup explicitly discloses the single-round limitation | Done (`README.md`) |
| 5 | Until step 3's hard gate exists, disclose per-GPU data sufficiency in every API/UI response rather than presenting uniform confidence | Done — `training_data_tier` field (`none`/`below_floor`/`sufficient`), 2026-07-11. Interim measure; does not replace step 3. |

---

## 4 · Excluded rows

| Category | Count | Reason |
|---|---:|---|
| CPU-inference (`gpu_name = N/A`) | 26 | Intel EMR and similar; no GPU to predict |
| NVIDIA Blackwell (B200, B300, GB200, GB300) | 206 | Explicitly out of v1 scope |
| NVIDIA RTX PRO 6000 Blackwell | 52 | Not in spec DB; out of scope |
| Heterogeneous multi-GPU | 20 | Mixed SKUs; cannot normalize to a single canonical ID |
| Intel Arc | 16 | Not in spec DB |
| Edge GPUs (Jetson) | 4 | Edge, not datacenter |
| GH200, H100-NVL, H100-PCIe, H200-NVL, L40S | 250 | Insufficient LLM rows or out of form-factor scope |
| **Total excluded** | **574** | |

Excluded rows are present in `mlperf_raw.parquet` with `gpu_in_model_scope` null or false.
They are **not removed** — exclusion happens at training time via the scope filter.

---

## 5 · Feature encoding decisions

### `precision` field

`precision` is 0 % populated across the entire corpus. MLPerf system names are hardware
identifiers (e.g., `8xMI300X_2xEPYC-9374F`), not precision-tagged. Regex extraction
returns `None` on all real submissions.

**`benchmark_accuracy_tier` is the reliable precision proxy** (100 % populated):

| Tier | Benchmark suffix | Typical precision regime |
|------|-----------------|--------------------------|
| `99.9` | `-99.9` | BF16 / FP16 (high-fidelity) |
| `99` | `-99` | FP8 or better |
| `base` | none | Throughput-optimized; any precision |

Distribution: `99` → 550 rows, `99.9` → 488 rows, `base` → 185 rows.

### `tokens_per_sample` (verified 2026-06-12)

These values are the divisor in `throughput_tok_per_sec_per_gpu`.

| Benchmark base | Value used | Source-verified | Delta | Source |
|---|---:|---:|---:|---|
| `gptj` | 128 | 128 (fixed) | 0.0 % | MLPerf benchmark spec |
| `llama2-70b` | 294 | 294.45 | < 0.2 % | `mlcommons/inference` language/llama2-70b/README.md |
| `mixtral-8x7b` | 145 | 144.84 | < 0.1 % | `mlcommons/inference` language/mixtral-8x7b README |
| `llama3.1-405b` | 294 | 294.45 | < 0.2 % | Same Open ORCA dataset as llama2-70b per MLPerf rules |

All values verified within ± 5 % tolerance. No parquet re-export required.

### `vram_gb` vs `gpu_vram_gb`

`vram_gb` (from the system JSON `accelerator_memory_capacity`) can reflect **total
system VRAM** for multi-GPU submissions. `gpu_vram_gb` (from `gpu_specs.yaml`) is
always per-GPU capacity.

**Rule: use `gpu_vram_gb` for the memory-fit constraint. Treat `vram_gb` as unreliable
for any system with `num_gpus > 1`.**

---

## 6 · Deduplication policy

No deduplication is applied. Rows from different rounds for the same GPU + benchmark
combination are intentional — they represent genuinely independent performance measurements
submitted to different MLPerf rounds.

**Primary key:** `(submitter, system_name, benchmark, scenario, round, division)`  
**Duplicate rows in corpus:** 0

---

## 7 · Known limitations

| ID | Limitation | Impact | Mitigation |
|----|-----------|--------|-----------|
| KL-01 | **MI355X v6.0-only** — single round, one benchmark family, 50 rows | AMD MAPE for MI355X has higher variance; no multi-round consistency check | Further AMD Dev Cloud calibration; 100-row gate. Not yet actionable — MI355X instances aren't available on the same Dev Cloud access used for MI300X. Interim: `training_data_tier: "below_floor"` discloses this per-request (§2 mitigation step 5). |
| KL-02 | `precision` field is 0 % populated | Cannot distinguish FP8 vs FP16 directly | Use `benchmark_accuracy_tier` as proxy |
| KL-03 | `tokens_per_sample` are rounded estimates | < 0.2 % systematic bias in `throughput_tok_per_sec_per_gpu` | Verified; rounding error is negligible |
| KL-04 | AMD corpus is 29 % of in-scope rows | AMD predictions are lower-confidence than NVIDIA | Relaxed AMD MAPE gate (< 20 % vs < 15 %) |
| KL-05 | No INVALID rows in current corpus | INVALID filter in training code is not currently exercised | Filter remains; future rounds may include INVALID |
| KL-06 | MI355X CDNA4 spec values are estimated | Roofline ceilings for MI355X may be slightly off | Reconfirm against AMD MI350 whitepaper before finalizing training data |
| KL-07 | `TOKENS_PER_SAMPLE` for llama3.1-405b inherits llama2-70b value | If Open ORCA output lengths differ across models, minor bias | Verify if llama3.1-405b round adds dataset-specific stats |
| KL-08 | Recommendation diversity is low: `mi355x` is the #1 Pareto pick in 73–80% of feasible model×tier queries, under all 4 ranking objectives | Undercuts the "workload-dependent, don't always pick the same GPU" framing this data supports the recommender for | Not a data-quality defect — mi355x measures as genuinely dominant at its listed price. Re-verify the $4.50/hr price assumption and/or add further Pareto axes in a future pass |