| ## Plan: A1 SI Baseline with Protocol C Evaluation and Expanded Parcels |
|
|
| Implement A1 as a frozen SI baseline with Protocol C cross-subject shared-space evaluation. Keep both 7 language ROIs and expanded Harvard-Oxford symmetric parcels, and report results by fold and run condition. |
|
|
| **Steps** |
| 1. Phase 0: Paper-first gate before coding |
| - Verify appendix details for ROI/parcellation and document any unresolved ambiguities. |
| - Lock assumptions and scope before implementation starts. |
|
|
| 2. Phase 1: Data manifest and integrity QC |
| - Build clean derivatives manifest with strict checks: folder and filename subject consistency, run whitelist 1-4, sidecar existence, unique subject-run keys. |
| - Save reject log and accepted manifest. |
| - Validate per-run TR and scan length from BIDS sidecars. |
|
|
| 3. Phase 2: Spatial harmonization and mask |
| - Resample BOLD to one canonical analysis grid. |
| - Build analysis mask and enforce left-right symmetry using hemisphere swap and mask intersection. |
| - Emit affine and shape and voxel-count QC report. |
|
|
| 4. Phase 3: Target definition with two parcellation tracks |
| - Track A: 7 language ROIs from sphere masks. |
| - Track B: Harvard-Oxford cortical maxprob-thr0-2mm with symmetric split for expanded parcel analysis. |
| - Resample labels with nearest-neighbor interpolation only. |
| - Build left-right parcel pairing map and parcel voxel index map. |
|
|
| 5. Phase 4: Annotation harmonization |
| - Build unified word-event table with run_id, condition, speaker_stream, word, onset_s, offset_s, duration_s, provenance. |
| - Apply approved fixed run mapping and mixed-run fallback flags. |
| - Validate monotonic and non-negative timing. |
| |
| 6. Phase 5: Frozen SI feature extraction |
| - Extract per-run, per-layer, per-word hidden-state features for approved models. |
| - Cache token-to-word aggregated embeddings and extraction metadata. |
| - No model fine-tuning in baseline phase. |
| |
| 7. Phase 6: Temporal alignment and normalization |
| - Convert word features to TR regressors with HRF convolution. |
| - Apply per-run z-score to regressors and BOLD targets. |
| - Enforce exact regressor-target TR length match after trimming policy. |
| |
| 8. Phase 7: Evaluation implementation (required) |
| - Protocol C: cross-subject shared-space split |
| - Use 21:3 subject split with 2-fold cross-validation over accepted subjects. |
| - Keep fold construction deterministic via a fixed seed. |
| - Compute metrics for both target sets: 7 ROIs and expanded parcels. |
| - Keep outputs clearly labeled by fold and run condition. |
| |
| 9. Phase 8: Outputs and reporting |
| - Save per-model, per-layer, per-subject, per-run, per-target, per-protocol tables. |
| - Save left-right asymmetry summaries and significance tables for parcel and ROI levels. |
| - Save Protocol C summary report across folds and run conditions. |
| |
| 10. Phase 9: Optional extension after baseline |
| - Add fine-tuning branch as a separate experiment only after frozen baseline is complete. |
| - Compare against frozen baseline outputs using matched protocols. |
| |
| **Relevant files** |
| - /home/mohith/ds005345/NeurIPS-2024-fmri-predictors-based-on-language-models-of-increasing-complexity-recover-brain-left-lateralization-Paper-Conference.pdf β paper and appendix reference gate. |
| - /home/mohith/ds005345/llms_brain_lateralization/README.md β reference pipeline order and 7 ROI context. |
| - /home/mohith/ds005345/llms_brain_lateralization/create_roi_masks.py β 7 language ROI coordinates and generation pattern. |
| - /home/mohith/ds005345/llms_brain_lateralization/analyze_results.ipynb β expanded Harvard-Oxford symmetric atlas workflow and nearest-neighbor resampling pattern. |
| - /home/mohith/ds005345/llms_brain_lateralization/extract_llm_activations.py β hidden-state extraction and token-to-word aggregation pattern. |
| - /home/mohith/ds005345/llms_brain_lateralization/fit_individual_subject.py β run-wise HRF regressor build and ridge CV template. |
| - /home/mohith/ds005345/llms_brain_lateralization/resample_fmri_data.py β canonical resampling pattern. |
| - /home/mohith/ds005345/llms_brain_lateralization/compute_mask.py β symmetric mask computation. |
| |
| **Verification** |
| 1. Paper gate checks |
| - Appendix ROI and parcel assumptions logged and aligned with implementation choices. |
| |
| 2. Spatial checks |
| - Atlas and BOLD match analysis affine and shape. |
| - Label interpolation nearest-neighbor only. |
| - Symmetric mask voxel counts and left-right balance logged. |
| |
| 3. Feature checks |
| - Hidden-state files complete for each run, model, and layer. |
| - Token-to-word aggregation diagnostics saved. |
| |
| 4. Modeling checks |
| - Frozen baseline confirmed with no weight updates. |
| - Protocol C fold table uses 21 train and 3 test subjects per fold with 2 folds. |
| - Each canonical run condition is evaluated per fold. |
| |
| 5. Reporting checks |
| - 7 ROI and expanded parcel outputs exist for both protocols. |
| - Side-by-side protocol comparison produced for every model. |
| |
| **Decisions** |
| - Baseline SI uses frozen inference plus ridge only. |
| - Protocol C cross-subject shared-space is the required baseline evaluation protocol. |
| - Random word-level 500/100 split is excluded from primary SI evaluation due temporal leakage risk after HRF convolution. |
| - Expanded ROI set uses Harvard-Oxford cortical maxprob-thr0-2mm with symmetric split. |
| - Fine-tuning remains a separate post-baseline experiment branch. |
| |
| ## Running On HF Jobs |
| |
| For this project, use HF Jobs rather than a Space. The A1 pipeline expects one root containing: |
| |
| - `project/` |
| - `data/` |
| - `derivatives/` |
| |
| The launcher in `project/hf_jobs.py` reconstructs that layout inside the HF container and then runs `project/code/run_a1_end_to_end.py` unchanged. |
|
|
| ### Required HF repos |
|
|
| Create two HF Dataset repos: |
|
|
| - one whose repo root mirrors the contents of local `data/` |
| - one whose repo root mirrors the contents of local `derivatives/` |
|
|
| Important: |
|
|
| - upload the folder contents, not the parent folder name |
| - the `data` repo should have files like `participants.tsv` and `sub-01/` at repo root |
| - the `derivatives` repo should have `sub-01/`, `sub-02/`, ... at repo root |
|
|
| ### Preview The Job |
|
|
| ```bash |
| cd /home/mohith/ds005345 |
| conda activate csai |
| python project/hf_jobs.py \ |
| --code-url https://huggingface.co/spaces/Mohith202/csai.git \ |
| --data-repo Mohith202/ds005345-data \ |
| --derivatives-repo Mohith202/ds005345-derivatives \ |
| --results-repo Mohith202/ds005345-a1-results \ |
| --secret-hf-token |
| ``` |
|
|
| This prints the generated job script and the `hf jobs run` command without launching it. |
|
|
| ### Launch The End-To-End Run |
|
|
| ```bash |
| cd /home/mohith/ds005345 |
| conda activate csai |
| python mohith/TEAM-9/hf_jobs.py \ |
| --code-url https://huggingface.co/spaces/Mohith202/csai.git \ |
| --data-repo Mohith202/ds005345-data \ |
| --derivatives-repo Mohith202/ds005345-derivatives \ |
| --results-repo Mohith202/ds005345-a1-results \ |
| --results-path hf_jobs/run_001 \ |
| --flavor a10g-small \ |
| --model-profile current \ |
| --protocols C \ |
| --secret-hf-token \ |
| --launch |
| ``` |
|
|
| ### Notes |
|
|
| - `--secret-hf-token` is recommended when the dataset repos are private. |
| - The job keeps the CUDA-enabled PyTorch already present in the container image and installs the rest of `project/requirements.txt` separately. |
| - By default bootstrap artifacts go to `outputs/a1_bootstrap_hf` inside the cloned repo. |
| - You can forward the normal end-to-end controls such as `--model-profile`, `--model-slug`, `--protocols`, `--reuse-caches`, `--skip-bootstrap`, `--skip-fit`, and `--skip-visualize`. |