# Spike 002a-mini-gpu-smoke — VERDICT **Status**: ✅ PASSED on local 5090 **Date**: 2026-05-26 **Wave**: 12 (closing the "zero GPU evidence" gap from cross-model review item #4) ## Headline `composer_replication` 3-channel loss composition runs cleanly on real GPU hardware. Qwen2.5-0.5B-Instruct on RTX 5090 sm_120 in bf16, 50 backward steps, loss 0.7354 → 0.00034 (99.95% reduction), all gradients finite throughout. Peak VRAM 5.31 GB (well under the ADR-001 8GB target). Median step time 480ms. ## Closes - The cross-model review's item #4: "Run Spike 002a-mini on the local 5090. ADR-001 made the choice; the spike was not run. Until then, the framework has zero GPU evidence of any kind." **Done.** - ADR-001's underlying claim that local 5090 is the right venue for this workload class. Verified: 50-step run completes in ~30 s wall-clock on the local 5090, vs an estimated 3-5 min cold-start cycle on Modal L4. - The "but the framework only runs on CPU" objection in V8. ## Acceptance criteria | Criterion | Target | Result | |---|---|---| | Model loads via `AutoModelForCausalLM` on `cuda` | bf16, no errors | ✅ 7.3 s | | 50 steps run end-to-end | No nan/inf | ✅ | | Loss decreases meaningfully | final < 50% × initial | ✅ final = 0.046% × initial | | Peak VRAM < 8 GB on 0.5B model | headroom check | ✅ 5.31 GB | | Step time stable | no thermal throttling, no swap | ✅ median 480ms, no outliers | | All gradients finite throughout | per-step finite check | ✅ | | sm_120 Blackwell architecture supported | not pre-Hopper-only | ✅ verified arch in `torch.cuda.get_arch_list()` | ## Per-channel behavior on GPU Same as CPU (Spike 006): LM-CE channel dominates, DPO channel contributes small nonzero gradient throughout, SDPO channel zero (shape-mismatch fallback — to exercise the SDPO channel on GPU, run with `align_sdpo_shapes` batch builder per Spike 006-strict's `test_sdpo_channel_actually_fires`). ## Memory profile | step | total | peak_mem_gb | step_time_ms | |------|-------|-------------|--------------| | 0 (post-warmup) | 0.7354 | 5.31 | ~500 | | 10 | 0.0067 | 5.31 | ~480 | | 25 | 0.0007 | 5.31 | ~480 | | 49 | 0.0003 | 5.31 | ~480 | Memory stays flat at 5.31 GB after warmup — no leak, no expanding activation buffers. (The 0.5B model in bf16 + Adam states + activations + DPO logit gradients all fit comfortably.) ## What this does NOT close - **Multi-replica / multi-process DiLoCo** (V2 partial gap). This spike is single-GPU. Real DiLoCo training across replicas is GPU-multi-process and not yet attempted. - **Composer-2.5-quality empirical results** (V8 partial gap). This spike verifies the framework runs on GPU; it does NOT verify the method improves model quality vs plain GRPO. That requires the full pipeline (real teacher rollouts + real GRPO rewards + a benchmark like SWE-bench-lite) and is the post-replication GPU phase ($30-100+). ## Files - `run_gpu_smoke.py` — 50-step GPU smoke runner with VRAM + step-time capture - `results/gpu_loss_curve.csv` — per-step metrics - `results/gpu_verdict.json` — programmatic verdict for CI/audit - `results/run.log` — actual successful run output ## Cost / time - $0 (local 5090, no Modal spend) - 35 s wall-clock total (7 s model load + 25 s training) - ~5 GB VRAM