Spaces:
Sleeping
Sleeping
| # GPU Perf Prophet — Data Card v0.1 | |
| > **Status:** Hand-authored baseline. This file will be superseded once `build_data_card()` | |
| > is implemented and `make data` is run. | |
| > §1, §2's MLPerf figures, and §4 are verified against `data/processed/mlperf_raw.parquet` | |
| > on 2026-06-12. §2 also covers the AMD Dev Cloud calibration rows added 2026-07-07, which | |
| > are a separate provenance layered on top of the MLPerf corpus at the training-features | |
| > stage (`data/processed/mlperf_features.parquet`) — they never enter `mlperf_raw.parquet`. | |
| --- | |
| ## 1 · Provenance | |
| | Field | Value | | |
| |-------|-------| | |
| | Primary source | MLCommons MLPerf Inference results repos v4.1, v5.0, v5.1, v6.0 | | |
| | Divisions | closed + open | | |
| | Ingestor | `src/data/mlperf_parser.py` | | |
| | GPU spec enrichment | `src/data/gpu_spec_db.py` → `data/gpu_specs.yaml` | | |
| | Processed artifact | `data/processed/mlperf_raw.parquet` | | |
| | Generated | 2026-06-12 | | |
| | Total parsed rows | **1,223** | | |
| | Unique submitters | 32 | | |
| | Unique system configs | 167 | | |
| | Result validity | 100% VALID (zero INVALID rows in corpus) | | |
| Source repos are sparse-checked-out git clones. Fetch scripts in `scripts/fetch_mlperf.sh`. | |
| Every parsed row carries `log_path` back to its originating `mlperf_log_summary.txt`. | |
| --- | |
| ## 2 · In-scope training corpus | |
| Rows entering model training must pass the `gpu_in_model_scope == True` filter | |
| (set in `data/gpu_specs.yaml`). CPU-only, heterogeneous, edge, and other | |
| out-of-scope GPUs are excluded. | |
| **Total MLPerf in-scope rows: 649** (2026-06-12). Plus 24 self-run AMD Dev Cloud MI300X | |
| calibration rows added 2026-07-07 — **total current in-scope rows: 673.** | |
| ### Per-GPU breakdown | |
| | Canonical ID | GPU | Vendor | Rounds present | MLPerf rows | Calibration rows | Total | Spec confidence | | |
| |---|---|---|---|---:|---:|---:|---| | |
| | `mi300x` | AMD Instinct MI300X | AMD | v4.1, v5.0, v5.1, v6.0 | 56 | 24 | **80** | verified | | |
| | `mi325x` | AMD Instinct MI325X | AMD | v5.0, v5.1, v6.0 | 82 | 0 | 82 | verified | | |
| | `mi355x` | AMD Instinct MI355X | AMD | **v6.0 only** | **50** | 0 | **50** | estimated | | |
| | `h100_sxm` | NVIDIA H100 SXM5 | NVIDIA | v4.1, v5.0, v5.1 | 178 | 0 | 178 | verified | | |
| | `h200_sxm` | NVIDIA H200 SXM | NVIDIA | v4.1, v5.0, v5.1 | 283 | 0 | 283 | verified | | |
| AMD total: **212 rows** (31 %) | |
| NVIDIA total: **461 rows** (69 %) | |
| The AMD/NVIDIA imbalance is the reason the AMD-specific MAPE target is relaxed to | |
| < 20 % (vs. < 15 % overall). | |
| **Calibration provenance:** the 24 MI300X rows were self-run on the AMD Developer Cloud | |
| (vLLM 0.23.0, ROCm 7.2.4) covering `gptj`, `llama2-70b`, `llama3.1-8b`, and `mixtral-8x7b` — | |
| `gptj` and `llama3.1-8b` had **zero** official MLPerf MI300X submissions before this. | |
| These rows measure a genuinely different regime from official submissions (no | |
| serving-stack tuning — see `README.md` § Recommendation accuracy); model-quality metrics | |
| reported anywhere in this project score against official MLPerf rows only, with | |
| calibration rows used purely as additional training signal, never as evaluation ground | |
| truth. | |
| ### Per-round in-scope row counts (MLPerf only; excludes calibration rows) | |
| | Round | mi300x | mi325x | mi355x | h100_sxm | h200_sxm | Round total | | |
| |-------|-------:|-------:|-------:|---------:|---------:|------------:| | |
| | v4.1 | 16 | — | — | 116 | 68 | 200 | | |
| | v5.0 | 4 | 16 | — | 54 | 156 | 230 | | |
| | v5.1 | 28 | 62 | — | 8 | 59 | 157 | | |
| | v6.0 | 8 | 4 | **50** | — | — | 62 | | |
| | **Total** | **56** | **82** | **50** | **178** | **283** | **649** | | |
| --- | |
| ## 3 · MI355X — v6.0-only GPU | |
| **MI355X is the only in-scope GPU with data from a single round.** | |
| This is the primary known limitation of the training corpus. | |
| ### What "v6.0-only" means | |
| | Property | Detail | | |
| |---|---| | |
| | First MLPerf round | v6.0 (no MI355X rows in v4.1, v5.0, or v5.1) | | |
| | In-scope training rows | 50 | | |
| | Benchmarks covered | `llama2-70b-99` (25 rows) and `llama2-70b-99.9` (25 rows) only | | |
| | Benchmarks *not* covered | `gptj`, `mixtral-8x7b`, `llama3.1-405b` | | |
| | Scenarios | Offline and Server | | |
| | Spec confidence | **estimated** — CDNA4 dense FLOPS derived from with-sparsity ÷ 2 | | |
| ### Why this matters for the model | |
| 1. **No round-over-round variance.** The model cannot check MI355X consistency across | |
| rounds, making it harder to detect submission anomalies or performance regressions. | |
| 2. **Single benchmark family.** The model sees MI355X only on `llama2-70b`. Predictions | |
| for MI355X on `gptj`, `mixtral-8x7b`, or `llama3.1-405b` are extrapolations, | |
| not interpolations. SHAP explanations will not decompose MI355X benchmark-level | |
| effects from GPU-level effects for those benchmarks. | |
| 3. **50 rows < 100-row minimum.** If MI355X does not meet the 100-row minimum gate | |
| after further AMD Dev Cloud calibration, it will be marked `enabled: false` in | |
| `gpu_specs.yaml` and excluded from v1 recommendations. The current value is | |
| `in_model_scope: true` as a placeholder pending that gate. **Until that gate is | |
| built, the API/UI discloses the shortfall per-request instead:** MI300X, MI325X, | |
| and MI355X (all under the 100-row floor) get `training_data_tier: "below_floor"` | |
| rather than being presented with the same confidence as a well-covered GPU — | |
| an interim, per-request disclosure, not a substitute for the hard gate above. | |
| 4. **Estimated spec values.** `gpu_specs.yaml` MI355X entries for `peak_tflops` are | |
| `sparse/2` derivations (flagged `spec_confidence: estimated`). These **must be | |
| reconfirmed** against the AMD MI350-series whitepaper before the training data | |
| is finalized. | |
| ### Mitigation plan | |
| | Step | Action | Status | | |
| |------|--------|--------| | |
| | 1 | Train with MI355X included; report MI355X MAPE separately; note in model card | Done | | |
| | 2 | Run MI300X + MI355X benchmarks on AMD Dev Cloud; add calibration rows | MI300X run 2026-07-07, but under its own ≥50-row / ≥3-LLM / ≥3-batch-size / 2-precision target: delivered 24 rows, 4 LLMs, 2 precisions, **zero batch-size variation**. Meets 2 of 4 axes. MI355X not yet run — Dev Cloud access covers MI300X instances only | | |
| | 3 | If MI355X row count ≥ 100 → keep `enabled: true`; else set `enabled: false` | Still blocked — MI355X remains at 50 rows | | |
| | 4 | Public writeup explicitly discloses the single-round limitation | Done (`README.md`) | | |
| | 5 | Until step 3's hard gate exists, disclose per-GPU data sufficiency in every API/UI response rather than presenting uniform confidence | Done — `training_data_tier` field (`none`/`below_floor`/`sufficient`), 2026-07-11. Interim measure; does not replace step 3. | | |
| --- | |
| ## 4 · Excluded rows | |
| | Category | Count | Reason | | |
| |---|---:|---| | |
| | CPU-inference (`gpu_name = N/A`) | 26 | Intel EMR and similar; no GPU to predict | | |
| | NVIDIA Blackwell (B200, B300, GB200, GB300) | 206 | Explicitly out of v1 scope | | |
| | NVIDIA RTX PRO 6000 Blackwell | 52 | Not in spec DB; out of scope | | |
| | Heterogeneous multi-GPU | 20 | Mixed SKUs; cannot normalize to a single canonical ID | | |
| | Intel Arc | 16 | Not in spec DB | | |
| | Edge GPUs (Jetson) | 4 | Edge, not datacenter | | |
| | GH200, H100-NVL, H100-PCIe, H200-NVL, L40S | 250 | Insufficient LLM rows or out of form-factor scope | | |
| | **Total excluded** | **574** | | | |
| Excluded rows are present in `mlperf_raw.parquet` with `gpu_in_model_scope` null or false. | |
| They are **not removed** — exclusion happens at training time via the scope filter. | |
| --- | |
| ## 5 · Feature encoding decisions | |
| ### `precision` field | |
| `precision` is 0 % populated across the entire corpus. MLPerf system names are hardware | |
| identifiers (e.g., `8xMI300X_2xEPYC-9374F`), not precision-tagged. Regex extraction | |
| returns `None` on all real submissions. | |
| **`benchmark_accuracy_tier` is the reliable precision proxy** (100 % populated): | |
| | Tier | Benchmark suffix | Typical precision regime | | |
| |------|-----------------|--------------------------| | |
| | `99.9` | `-99.9` | BF16 / FP16 (high-fidelity) | | |
| | `99` | `-99` | FP8 or better | | |
| | `base` | none | Throughput-optimized; any precision | | |
| Distribution: `99` → 550 rows, `99.9` → 488 rows, `base` → 185 rows. | |
| ### `tokens_per_sample` (verified 2026-06-12) | |
| These values are the divisor in `throughput_tok_per_sec_per_gpu`. | |
| | Benchmark base | Value used | Source-verified | Delta | Source | | |
| |---|---:|---:|---:|---| | |
| | `gptj` | 128 | 128 (fixed) | 0.0 % | MLPerf benchmark spec | | |
| | `llama2-70b` | 294 | 294.45 | < 0.2 % | `mlcommons/inference` language/llama2-70b/README.md | | |
| | `mixtral-8x7b` | 145 | 144.84 | < 0.1 % | `mlcommons/inference` language/mixtral-8x7b README | | |
| | `llama3.1-405b` | 294 | 294.45 | < 0.2 % | Same Open ORCA dataset as llama2-70b per MLPerf rules | | |
| All values verified within ± 5 % tolerance. No parquet re-export required. | |
| ### `vram_gb` vs `gpu_vram_gb` | |
| `vram_gb` (from the system JSON `accelerator_memory_capacity`) can reflect **total | |
| system VRAM** for multi-GPU submissions. `gpu_vram_gb` (from `gpu_specs.yaml`) is | |
| always per-GPU capacity. | |
| **Rule: use `gpu_vram_gb` for the memory-fit constraint. Treat `vram_gb` as unreliable | |
| for any system with `num_gpus > 1`.** | |
| --- | |
| ## 6 · Deduplication policy | |
| No deduplication is applied. Rows from different rounds for the same GPU + benchmark | |
| combination are intentional — they represent genuinely independent performance measurements | |
| submitted to different MLPerf rounds. | |
| **Primary key:** `(submitter, system_name, benchmark, scenario, round, division)` | |
| **Duplicate rows in corpus:** 0 | |
| --- | |
| ## 7 · Known limitations | |
| | ID | Limitation | Impact | Mitigation | | |
| |----|-----------|--------|-----------| | |
| | KL-01 | **MI355X v6.0-only** — single round, one benchmark family, 50 rows | AMD MAPE for MI355X has higher variance; no multi-round consistency check | Further AMD Dev Cloud calibration; 100-row gate. Not yet actionable — MI355X instances aren't available on the same Dev Cloud access used for MI300X. Interim: `training_data_tier: "below_floor"` discloses this per-request (§2 mitigation step 5). | | |
| | KL-02 | `precision` field is 0 % populated | Cannot distinguish FP8 vs FP16 directly | Use `benchmark_accuracy_tier` as proxy | | |
| | KL-03 | `tokens_per_sample` are rounded estimates | < 0.2 % systematic bias in `throughput_tok_per_sec_per_gpu` | Verified; rounding error is negligible | | |
| | KL-04 | AMD corpus is 29 % of in-scope rows | AMD predictions are lower-confidence than NVIDIA | Relaxed AMD MAPE gate (< 20 % vs < 15 %) | | |
| | KL-05 | No INVALID rows in current corpus | INVALID filter in training code is not currently exercised | Filter remains; future rounds may include INVALID | | |
| | KL-06 | MI355X CDNA4 spec values are estimated | Roofline ceilings for MI355X may be slightly off | Reconfirm against AMD MI350 whitepaper before finalizing training data | | |
| | KL-07 | `TOKENS_PER_SAMPLE` for llama3.1-405b inherits llama2-70b value | If Open ORCA output lengths differ across models, minor bias | Verify if llama3.1-405b round adds dataset-specific stats | | |
| | KL-08 | Recommendation diversity is low: `mi355x` is the #1 Pareto pick in 73–80% of feasible model×tier queries, under all 4 ranking objectives | Undercuts the "workload-dependent, don't always pick the same GPU" framing this data supports the recommender for | Not a data-quality defect — mi355x measures as genuinely dominant at its listed price. Re-verify the $4.50/hr price assumption and/or add further Pareto axes in a future pass | | |