Buckets:

glennmatlin's picture
|
download
raw
11.2 kB

SLURM GPU/Walltime Catalog (Phase A.2 of compute audit)

Authoritative catalog of every SLURM script in scripts/slurm/, extracted by grep of #SBATCH headers on 2026-05-18. Source of truth for "which GPU per phase" claims in sections/appendix/11-compute.tex.

PACE partition → GPU mapping (per TDA-91, TDA-99)

Partition GPU types accepted Use
coe-gpu H200 144 GB, H100 80 GB (predominantly) Production attribution + unlearning + evaluation
ice-gpu A100 / L40S / V100 mixed Some attribution + evaluation fallback
ice-bw-gpu RTX PRO 6000 BW 96 GB (Blackwell), older RTX Enrichment fallback for low-VRAM
coc-gpu (older partition; A100-heavy) Early attribution scripts
ice-cpu CPU-only Dot-product scoring, preconditioner merge, aggregation
RunPod / Modal external Some unlearning runs (not in slurm logs)

Phase 1: Taxonomy Enrichment (TDA-91)

Script Partition GPU type / constraint Walltime Notes
enrichment/enrich_sidecar_gpu.sbatch ice-gpu, coe-gpu, ice-bw-gpu nvidia-gpu (any NVIDIA, auto-detected) 16h PRODUCTION: per-shard auto-detects H200/H100/RTX PRO 6000/L40S/V100/RTX 6000; 2,826 array tasks × 20 shards. Excludes bad node atl1-1-03-014-16-0.
enrichment/enrich_pool_gpu.sbatch (auto) `HX00 L40S A100-80GB
enrichment/enrich_pool_150B_gpu.sh (auto) gpu:1 (any) 16h 150B pool refinement
eda/sidecar_eda_array_phoenix.sbatch (PACE Phoenix) varies Sidecar EDA validation passes

Per-shard .stats.json files capture elapsed_seconds per task (TDA-122 instrumentation). Uploaded to Cloudflare R2 — bucket soc127-dedup, prefixes soc91-labels/*.stats.json (SLURM) and soc91-stats/... (Modal). The schema (src/dolma/sidecar_writer.py finalize_stats()) is {docs_classified, elapsed_seconds, docs_per_second, shard, source_family, docs_failed}does NOT include gpu_type. GPU-type breakdown requires PACE SLURM .out logs (if available).

Phase 2a: Attribution — Training Gradient Index (TDA-156, TDA-116)

Script Partition GPU Walltime Notes
attribution/tracstar_build.sbatch (auto) gpu:1 6h Training index build — Mode A canonical
attribution/tracstar_reduce.sbatch (auto) gpu:1 12h Reduce (mean query gradient per benchmark)
attribution/data_index.sh (auto) gpu:v100:1 4h V100 test script (early, abandoned)

Per TDA-156: actual run used 316 shards × 5.68M docs. Mode A 8,824 GPU-hr in tab-mode-comparison.tex is projection for full 39.7M corpus, not the 5.68M actual run.

Phase 2b: Attribution — Query Gradient Indexes (TDA-170 BBH + earlier)

Script Partition GPU Walltime Notes
attribution/query_index.sh (auto) gpu:a100:1 12h Query reduce for GSM8K/SocialIQA/MMLU SS/MMLU STEM (Mode B / Mode A query phase)
attribution/query_index_instruct.sh coe-gpu gpu:1 16h Instruct-model query indexes (TDA-168)
attribution/olmes_query_build.sh coc-gpu gpu:1 3h OLMES-format query manifest build
attribution/tracstar_query_build.sbatch (auto) gpu:1 1h TrackStar query gradient build

Per TDA-170: 6 jobs × 3-5 min on H100 ≈ 0.5 GPU-hr total per benchmark variant. Across 6 benchmarks × 2 variants (base+instruct) ≈ 10 GPU-hr aggregated.

Phase 2c: Preconditioner Construction (TDA-152 base, TDA-168 instruct)

Script Partition GPU Walltime Notes
attribution/trackstar_preconditioner.sbatch (auto) gpu:8 (FSDP, H100 per commit 50fc0cd) 12h Preconditioner build — 100K sample
attribution/precond_value.sbatch (auto) (gpu inherited) 2h Value preconditioner
attribution/precond_query.sbatch (auto) (gpu inherited) 2h Query preconditioner
attribution/precond_chunk.sbatch (auto) (gpu inherited) 2h Chunked preconditioner
attribution/precond_mix.sbatch ice-cpu — (CPU) 30 min Combination step (CPU)
attribution/precond_merge_and_query.sbatch ice-cpu — (CPU) 1h Merge + query (CPU)
attribution/precond_cache_data.sbatch ice-cpu — (CPU) 2h Data caching (CPU)
attribution/precond_upload.sbatch ice-cpu — (CPU) 1h Upload (CPU)
attribution/launch_preconditioner.sh — (launcher) Orchestrator

Estimate: 8 H100 × 12h × 2 builds (base + instruct) = 192 GPU-hr. Plus CPU-side merge/mix/cache/upload (negligible GPU-hr).

Phase 2d: Attribution — Production Dot-Product Scoring (TDA-170, TDA-171)

ALL CPU, not GPU:

Script Partition GPU Walltime Notes
attribution/tracstar_dot_score.sbatch (default CPU) none 2h Per-shard dot-product against training index
attribution/tracstar_dot_score_array.sbatch (default CPU) none 1h Array variant
attribution/tracstar_score.sbatch (auto) gpu:1 6h (alternative: GPU-based score — only used for some early runs)
attribution/tracstar_aggregate.sbatch (default CPU) none 1h Aggregation
attribution/tracstar_bin_aggregate.sbatch (default CPU) none 6h Bin-level aggregation
attribution/tracstar_bin_aggregate_perquery.sbatch (default CPU) none 6h Per-query bin aggregation
attribution/tracstar_bin_aggregate_split.sbatch (default CPU) none 12h Correctness-split aggregation
attribution/tracstar_top2k_extract.sbatch ice-cpu none 2h Top-2k document extraction
attribution/score.sh ice-gpu gpu:a100:1 12h Alternative score path (some early runs)
attribution/influence.sh coc-gpu gpu:a100:1 4h Older influence script
attribution/influence_simple.sh ice-gpu gpu:a100:1 4h Smoke test

Per TDA-170: production scoring is 128 CPU array tasks per benchmark × 4 min/worker ≈ 8 CPU-hours. Should not be counted as GPU-hours in §L.2.

Phase 3: Unlearning (TDA-103, TDA-106, TDA-145, TDA-143, TDA-154)

Script Partition GPU Walltime Notes
unlearning/unlearn_ngdiff.sh coe-gpu gpu:1 (H200 per commit cf94fc9) 16h Main NGDiff trainer
unlearning/run_single_topics.sh (orchestrator) (via unlearn_ngdiff) 24 topics × 3 chained jobs = 72 slurm jobs
unlearning/run_cross_topics.sh (orchestrator) (via unlearn_ngdiff) Multi-bin / cross-topic
unlearning/chain_unlearn.sh (orchestrator) (via unlearn_ngdiff) 3-job chain with auto-resume
unlearning/sweep_ngdiff.sh (orchestrator) (via unlearn_ngdiff) Hyperparameter sweep: 5 topics × forget×retain×LR matrix (TDA-145)
unlearning/sweep_benchmark.sh (orchestrator) Benchmark sweep
unlearning/run_single_topics_runpod.sh RunPod (not SLURM) unlimited 24 topics sequentially on RunPod node — invisible to slurm logs
unlearning/eval_unlearn.sh coe-gpu gpu:1 3h Per-checkpoint eval (every 250 steps × 4 benchmarks)
unlearning/eval_fast.sh coe-gpu gpu:1 1.5h Fast eval (200 samples/task)
unlearning/eval_baseline.sh coe-gpu gpu:1 4h Baseline reference eval
unlearning/eval_wikitext_baseline.sh coe-gpu gpu:1 1h Wikitext baseline
unlearning/forget_ppl.sh coe-gpu gpu:1 1.5h Forget-set perplexity check
unlearning/build_6t_cache.sh coe-gpu (no gres line, but partition is GPU) 2h Data cache prep
unlearning/build_hf_cache.sh coe-gpu (no gres) 6h HF cache prep

Floor estimate (slurm-only):

  • 24 topics × 3 chained jobs × 16h = 1,152 GPU-hr for single-topic influence-guided
    • 24 topics × 3 × 16h = 1,152 GPU-hr for single-topic random baseline
    • multi-bin runs (TDA-143) ≈ several hundred GPU-hr
    • hyperparameter sweep (sweep_ngdiff.sh, TDA-145) ≈ several hundred GPU-hr
    • eval passes (eval_unlearn / eval_fast every 250 steps × 4 benchmarks): substantial
    • RunPod runs (NOT in slurm; need separate accounting)

Conservative slurm-only floor: ~3,000 GPU-hr on H200 (coe-gpu). With RunPod + multi-bin + sweep + evals: likely 4,000-6,000 GPU-hr H200-equivalent.

The current manuscript claim of "~2,000 GPU-hours on A100 80GB" is wrong on both axes: hardware (H200 not A100) and magnitude (under-counts by 1.5-3×).

Phase 4: Evaluation (post-unlearning OLMES)

Script Partition GPU Walltime Notes
evaluation/eval_arc.sh coe-gpu, ice-gpu gpu:1 4h ARC-Easy / ARC-Challenge eval
evaluation/eval_bbh.sh coe-gpu, ice-gpu gpu:1 2h BBH eval (post-unlearning + accuracy baseline)
evaluation/olmes_instruct_base.sbatch coe-gpu gpu:1 16h OLMES full eval batch
evaluation/olmes_instruct_base_n1.sbatch coe-gpu gpu:1 2h OLMES single benchmark

Predominantly H200/H100 (coe-gpu). ice-gpu fallback (A100/L40S) for capacity.

Phase 5: Dedup / Sampling / Manifest (TDA-90, TDA-95)

Script Partition GPU Walltime Notes
dedup/*.sbatch varies mostly CPU varies Bloom filter dedup, parquet/jsonl transforms
sampling/extract_stratified_docs.sbatch (CPU) none Stratified sampling
manifest/build_corpus_manifest.sh (CPU) none Manifest build

Not GPU-bound; these CPU phases are not counted as GPU-hr.

Aggregated raw-walltime ESTIMATE (slurm-only, awaiting log walk)

Phase GPU type Estimated raw GPU-hr H200-equiv factor H200-equiv GPU-hr
1 Enrichment mixed (H200/H100/RTX PRO 6000/L40S/V100/RTX 6000) 8,000-12,000 (mixed; ~0.7-1.0 avg) 7,000-10,000
2a Training index H100 80GB 1,200-2,000 1.00 1,200-2,000
2b Query indexes H100/A100 10-50 ~0.7 7-35
2c Preconditioner 8× H100 (2 builds) 192 1.00 192
2d Calibration/smoke H100/A100 400-600 ~0.7 280-420
3 Unlearning H200 + RunPod 3,000-6,000 1.00 (H200) / ~0.32 (A100) 3,000-6,000
4 Evaluation H200/H100 mostly 200-500 1.00 200-500
Total ~13,000-21,000 raw ~12,000-19,000 H200-equiv

Floor (sum of mid-range estimates): ~16,000 H200-equiv GPU-hr. Ceiling (sum of high-range): ~24,000 H200-equiv GPU-hr.

The current manuscript headline of 30,000 mixed-tier may actually be roughly correct as a RAW total, but the H200-equivalent is lower because some enrichment hours were on slower GPUs. However, this estimate is slurm-only — adding RunPod / Modal Labs work + failed runs + sweep work could push the real total higher.

Available autonomously: R2 enrichment .stats.json sidecars (raw GPU-hr only, no gpu_type) + HF bucket timestamps for Phase 2/4 wall-clock floor. Still needs human input: RunPod billing CSV, Modal app history, confirmation that 156.1 exploratory ran on H100 vs A100.

Xet Storage Details

Size:
11.2 kB
·
Xet hash:
7d01c5a5c313165b45d52df5a0c4ac5dddfa681d7230b7644f2b2be22c5115b5

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.