File size: 1,050 Bytes
141bacd | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 | # scripts/pan_skin
canonical end-to-end pipeline for the pan-skin corpus. steps are numbered in the
order they need to run. everything writes into `data/corpus/pan_skin/`,
`checkpoints/pan_skin/`, and `figures/`.
| step | script | what it does |
|---|---|---|
| 01-03 | `01_download_tier_a.sh` .. `03_download_tier_c.sh` | fetch tier A/B/C raw h5ads + counts |
| 06 | `06_build_per_dataset_h5ads.py` | per-dataset QC + h5ad build |
| 07 | `07_build_shared_hvgs_and_pca.py` | shared HVGs + PCA basis over the union |
| 08 | `08_assign_labels.py` (+ `08b_curated_label_override.py`) | marker-scored labels, Sulic overrides |
| 10 | `10_build_corpus.py` | merged harmonized corpus.h5ad |
| 20 | `20_train_panda.py` | PANDA-MLP with hybrid sampling |
| 30 | `30_zero_shot_aldrich.py` | inference on Aldrich / Dingwall (held-out target) |
| 40 | `40_heldout_5fold_cv.py` | 5-fold CV, PANDA retrained from scratch per fold |
`run_all.sh` chains these in order. marker panels: `known_skin_tfs.yaml`; ontology
map: `ontology_map.yaml`.
|