Spaces:
Sleeping
GPU Perf Prophet — Data Card v0.1
Status: Hand-authored baseline. This file will be superseded once
build_data_card()is implemented andmake datais run. §1, §2's MLPerf figures, and §4 are verified againstdata/processed/mlperf_raw.parqueton 2026-06-12. §2 also covers the AMD Dev Cloud calibration rows added 2026-07-07, which are a separate provenance layered on top of the MLPerf corpus at the training-features stage (data/processed/mlperf_features.parquet) — they never entermlperf_raw.parquet.
1 · Provenance
| Field | Value |
|---|---|
| Primary source | MLCommons MLPerf Inference results repos v4.1, v5.0, v5.1, v6.0 |
| Divisions | closed + open |
| Ingestor | src/data/mlperf_parser.py |
| GPU spec enrichment | src/data/gpu_spec_db.py → data/gpu_specs.yaml |
| Processed artifact | data/processed/mlperf_raw.parquet |
| Generated | 2026-06-12 |
| Total parsed rows | 1,223 |
| Unique submitters | 32 |
| Unique system configs | 167 |
| Result validity | 100% VALID (zero INVALID rows in corpus) |
Source repos are sparse-checked-out git clones. Fetch scripts in scripts/fetch_mlperf.sh.
Every parsed row carries log_path back to its originating mlperf_log_summary.txt.
2 · In-scope training corpus
Rows entering model training must pass the gpu_in_model_scope == True filter
(set in data/gpu_specs.yaml). CPU-only, heterogeneous, edge, and other
out-of-scope GPUs are excluded.
Total MLPerf in-scope rows: 649 (2026-06-12). Plus 24 self-run AMD Dev Cloud MI300X calibration rows added 2026-07-07 — total current in-scope rows: 673.
Per-GPU breakdown
| Canonical ID | GPU | Vendor | Rounds present | MLPerf rows | Calibration rows | Total | Spec confidence |
|---|---|---|---|---|---|---|---|
mi300x |
AMD Instinct MI300X | AMD | v4.1, v5.0, v5.1, v6.0 | 56 | 24 | 80 | verified |
mi325x |
AMD Instinct MI325X | AMD | v5.0, v5.1, v6.0 | 82 | 0 | 82 | verified |
mi355x |
AMD Instinct MI355X | AMD | v6.0 only | 50 | 0 | 50 | estimated |
h100_sxm |
NVIDIA H100 SXM5 | NVIDIA | v4.1, v5.0, v5.1 | 178 | 0 | 178 | verified |
h200_sxm |
NVIDIA H200 SXM | NVIDIA | v4.1, v5.0, v5.1 | 283 | 0 | 283 | verified |
AMD total: 212 rows (31 %)
NVIDIA total: 461 rows (69 %)
The AMD/NVIDIA imbalance is the reason the AMD-specific MAPE target is relaxed to < 20 % (vs. < 15 % overall).
Calibration provenance: the 24 MI300X rows were self-run on the AMD Developer Cloud
(vLLM 0.23.0, ROCm 7.2.4) covering gptj, llama2-70b, llama3.1-8b, and mixtral-8x7b —
gptj and llama3.1-8b had zero official MLPerf MI300X submissions before this.
These rows measure a genuinely different regime from official submissions (no
serving-stack tuning — see README.md § Recommendation accuracy); model-quality metrics
reported anywhere in this project score against official MLPerf rows only, with
calibration rows used purely as additional training signal, never as evaluation ground
truth.
Per-round in-scope row counts (MLPerf only; excludes calibration rows)
| Round | mi300x | mi325x | mi355x | h100_sxm | h200_sxm | Round total |
|---|---|---|---|---|---|---|
| v4.1 | 16 | — | — | 116 | 68 | 200 |
| v5.0 | 4 | 16 | — | 54 | 156 | 230 |
| v5.1 | 28 | 62 | — | 8 | 59 | 157 |
| v6.0 | 8 | 4 | 50 | — | — | 62 |
| Total | 56 | 82 | 50 | 178 | 283 | 649 |
3 · MI355X — v6.0-only GPU
MI355X is the only in-scope GPU with data from a single round.
This is the primary known limitation of the training corpus.
What "v6.0-only" means
| Property | Detail |
|---|---|
| First MLPerf round | v6.0 (no MI355X rows in v4.1, v5.0, or v5.1) |
| In-scope training rows | 50 |
| Benchmarks covered | llama2-70b-99 (25 rows) and llama2-70b-99.9 (25 rows) only |
| Benchmarks not covered | gptj, mixtral-8x7b, llama3.1-405b |
| Scenarios | Offline and Server |
| Spec confidence | estimated — CDNA4 dense FLOPS derived from with-sparsity ÷ 2 |
Why this matters for the model
No round-over-round variance. The model cannot check MI355X consistency across rounds, making it harder to detect submission anomalies or performance regressions.
Single benchmark family. The model sees MI355X only on
llama2-70b. Predictions for MI355X ongptj,mixtral-8x7b, orllama3.1-405bare extrapolations, not interpolations. SHAP explanations will not decompose MI355X benchmark-level effects from GPU-level effects for those benchmarks.50 rows < 100-row minimum. If MI355X does not meet the 100-row minimum gate after further AMD Dev Cloud calibration, it will be marked
enabled: falseingpu_specs.yamland excluded from v1 recommendations. The current value isin_model_scope: trueas a placeholder pending that gate. Until that gate is built, the API/UI discloses the shortfall per-request instead: MI300X, MI325X, and MI355X (all under the 100-row floor) gettraining_data_tier: "below_floor"rather than being presented with the same confidence as a well-covered GPU — an interim, per-request disclosure, not a substitute for the hard gate above.Estimated spec values.
gpu_specs.yamlMI355X entries forpeak_tflopsaresparse/2derivations (flaggedspec_confidence: estimated). These must be reconfirmed against the AMD MI350-series whitepaper before the training data is finalized.
Mitigation plan
| Step | Action | Status |
|---|---|---|
| 1 | Train with MI355X included; report MI355X MAPE separately; note in model card | Done |
| 2 | Run MI300X + MI355X benchmarks on AMD Dev Cloud; add calibration rows | MI300X run 2026-07-07, but under its own ≥50-row / ≥3-LLM / ≥3-batch-size / 2-precision target: delivered 24 rows, 4 LLMs, 2 precisions, zero batch-size variation. Meets 2 of 4 axes. MI355X not yet run — Dev Cloud access covers MI300X instances only |
| 3 | If MI355X row count ≥ 100 → keep enabled: true; else set enabled: false |
Still blocked — MI355X remains at 50 rows |
| 4 | Public writeup explicitly discloses the single-round limitation | Done (README.md) |
| 5 | Until step 3's hard gate exists, disclose per-GPU data sufficiency in every API/UI response rather than presenting uniform confidence | Done — training_data_tier field (none/below_floor/sufficient), 2026-07-11. Interim measure; does not replace step 3. |
4 · Excluded rows
| Category | Count | Reason |
|---|---|---|
CPU-inference (gpu_name = N/A) |
26 | Intel EMR and similar; no GPU to predict |
| NVIDIA Blackwell (B200, B300, GB200, GB300) | 206 | Explicitly out of v1 scope |
| NVIDIA RTX PRO 6000 Blackwell | 52 | Not in spec DB; out of scope |
| Heterogeneous multi-GPU | 20 | Mixed SKUs; cannot normalize to a single canonical ID |
| Intel Arc | 16 | Not in spec DB |
| Edge GPUs (Jetson) | 4 | Edge, not datacenter |
| GH200, H100-NVL, H100-PCIe, H200-NVL, L40S | 250 | Insufficient LLM rows or out of form-factor scope |
| Total excluded | 574 |
Excluded rows are present in mlperf_raw.parquet with gpu_in_model_scope null or false.
They are not removed — exclusion happens at training time via the scope filter.
5 · Feature encoding decisions
precision field
precision is 0 % populated across the entire corpus. MLPerf system names are hardware
identifiers (e.g., 8xMI300X_2xEPYC-9374F), not precision-tagged. Regex extraction
returns None on all real submissions.
benchmark_accuracy_tier is the reliable precision proxy (100 % populated):
| Tier | Benchmark suffix | Typical precision regime |
|---|---|---|
99.9 |
-99.9 |
BF16 / FP16 (high-fidelity) |
99 |
-99 |
FP8 or better |
base |
none | Throughput-optimized; any precision |
Distribution: 99 → 550 rows, 99.9 → 488 rows, base → 185 rows.
tokens_per_sample (verified 2026-06-12)
These values are the divisor in throughput_tok_per_sec_per_gpu.
| Benchmark base | Value used | Source-verified | Delta | Source |
|---|---|---|---|---|
gptj |
128 | 128 (fixed) | 0.0 % | MLPerf benchmark spec |
llama2-70b |
294 | 294.45 | < 0.2 % | mlcommons/inference language/llama2-70b/README.md |
mixtral-8x7b |
145 | 144.84 | < 0.1 % | mlcommons/inference language/mixtral-8x7b README |
llama3.1-405b |
294 | 294.45 | < 0.2 % | Same Open ORCA dataset as llama2-70b per MLPerf rules |
All values verified within ± 5 % tolerance. No parquet re-export required.
vram_gb vs gpu_vram_gb
vram_gb (from the system JSON accelerator_memory_capacity) can reflect total
system VRAM for multi-GPU submissions. gpu_vram_gb (from gpu_specs.yaml) is
always per-GPU capacity.
Rule: use gpu_vram_gb for the memory-fit constraint. Treat vram_gb as unreliable
for any system with num_gpus > 1.
6 · Deduplication policy
No deduplication is applied. Rows from different rounds for the same GPU + benchmark combination are intentional — they represent genuinely independent performance measurements submitted to different MLPerf rounds.
Primary key: (submitter, system_name, benchmark, scenario, round, division)
Duplicate rows in corpus: 0
7 · Known limitations
| ID | Limitation | Impact | Mitigation |
|---|---|---|---|
| KL-01 | MI355X v6.0-only — single round, one benchmark family, 50 rows | AMD MAPE for MI355X has higher variance; no multi-round consistency check | Further AMD Dev Cloud calibration; 100-row gate. Not yet actionable — MI355X instances aren't available on the same Dev Cloud access used for MI300X. Interim: training_data_tier: "below_floor" discloses this per-request (§2 mitigation step 5). |
| KL-02 | precision field is 0 % populated |
Cannot distinguish FP8 vs FP16 directly | Use benchmark_accuracy_tier as proxy |
| KL-03 | tokens_per_sample are rounded estimates |
< 0.2 % systematic bias in throughput_tok_per_sec_per_gpu |
Verified; rounding error is negligible |
| KL-04 | AMD corpus is 29 % of in-scope rows | AMD predictions are lower-confidence than NVIDIA | Relaxed AMD MAPE gate (< 20 % vs < 15 %) |
| KL-05 | No INVALID rows in current corpus | INVALID filter in training code is not currently exercised | Filter remains; future rounds may include INVALID |
| KL-06 | MI355X CDNA4 spec values are estimated | Roofline ceilings for MI355X may be slightly off | Reconfirm against AMD MI350 whitepaper before finalizing training data |
| KL-07 | TOKENS_PER_SAMPLE for llama3.1-405b inherits llama2-70b value |
If Open ORCA output lengths differ across models, minor bias | Verify if llama3.1-405b round adds dataset-specific stats |
| KL-08 | Recommendation diversity is low: mi355x is the #1 Pareto pick in 73–80% of feasible model×tier queries, under all 4 ranking objectives |
Undercuts the "workload-dependent, don't always pick the same GPU" framing this data supports the recommender for | Not a data-quality defect — mi355x measures as genuinely dominant at its listed price. Re-verify the $4.50/hr price assumption and/or add further Pareto axes in a future pass |