Amphora_NeuroText / README.md
ffh92r32rm0's picture
Update README.md
10f940f verified
|
Raw
History Blame Contribute Delete
5.43 kB
metadata
license: mit
tags:
  - neuroscience
  - fMRI
  - brain-decoding
  - text-to-brain
  - audio-to-brain
  - naturalistic-fMRI
  - ROI-prediction
  - cognitive-neuroscience
  - brain-encoding
  - zero-shot
  - whisper
  - Qwen4B
library_name: custom
language:
  - en
datasets:
  - narratives
  - little-prince
  - hcp
  - cneuromod
pipeline_tag: feature-extraction

Amphora NeuroText v4 β€” Audio & Text β†’ Brain ROI Activation

Small MLP brain encoder: predict which brain regions activate in response to any audio or text β€” no fMRI required at inference time.

Open In Colab

Trained exclusively on real naturalistic fMRI data.

This is a demo, not representative of the full model**.

Predicts activation across 56 canonical brain ROIs (HCP MMP1.0 parcellation).


Headline Result

Audio model (Whisper v4) beats TRIBE v2 β€” the Meta AI model that won the Algonauts 2025 competition β€” by +4.2% on a 23-subject cross-subject holdout:

Model Held-out R Notes
NeuroText Whisper v4 0.257 Single shared model, no per-subject fine-tuning
TRIBE v2 (Meta AI) 0.215 Published; per-subject fine-tuned; video+audio+text

All 7 brain networks exceeded TRIBE v2 at a fraction of the compute. Biggest leads: Frontal (+0.209), Default Mode (+0.221), Subcortical (+0.234). These predictions are made on a coarser 56 ROI level, future work will be devoted towards increasing prediction fidelity.


Models

File Input Val R Holdout R Notes
text2roi_whisper_v4.pt Audio (Whisper-large-v3, 1280d) 0.217 0.257 Recommended for audio. Beats TRIBE v2 (+4.2%)
text2roi_combined_v4.pt Text+Audio (3840d, modality dropout) 0.192 β€” Text or combined inference; recommended for text
text2roi_qwen3_v8.pt Text (Qwen3-Embedding-4B, 2560d) 0.115 β€” Text-only; cross-dataset generalization improving

Val R = mean Pearson R across 56 ROIs on held-out subjects (honest β€” see below). Holdout R = independent eval on 23 never-seen subjects.

Note on previous models (v2/v3): Models in this repo before July 2026 reported inflated val_R values (0.239–0.413) due to within-subject train/val splits. Those numbers are not comparable to v4. The v4 models use honest per-subject z-scoring and per-subject holdout splits.


Why These Numbers Are Honest

Previous NeuroText versions had inflated val_R from a within-subject split: the model saw the same subjects in both train and val, and learned subject-level baseline BOLD activations. Cross-subject, that memorized baseline is useless.

v4 fixes both issues:

  1. Per-subject z-scoring β€” each subject's ROI activations are z-scored independently before training, so the model learns stimulus-driven fluctuations rather than who each subject is.
  2. Per-subject holdout split β€” 15% of subjects per dataset are excluded from training entirely. Zero subject overlap guaranteed.
  3. Single shared model β€” no per-subject adaptation at inference. TRIBE requires fine-tuning on each test subject; Amphora does not.

Audio Model β€” Top ROIs (held-out eval)

ROI R vs TRIBE
ACC 0.438 +0.368
STG 0.422 +0.302
Thalamus 0.408 +0.358
V1 0.393 +0.073
LP_R 0.380 +0.290
mPFC_dmn 0.379 +0.299
AI 0.343 +0.253
HPC_L 0.337 +0.247
V2 0.336 +0.036
dACC 0.330 +0.250

53/56 ROIs R>0.10 Β· 41/56 R>0.20 Β· 20/56 R>0.30


Quick Start

from huggingface_hub import hf_hub_download
from predict import predict_audio, predict_text, top_rois

REPO = "ffh92r32rm0/Amphora_NeuroText"

# Audio β†’ brain regions (recommended)
ckpt = hf_hub_download(REPO, "text2roi_whisper_v4.pt")
roi_map = predict_audio("clip.wav", ckpt)
print(top_rois(roi_map, n=5))
# β†’ [('ACC', 0.44), ('STG', 0.42), ('Thalamus', 0.41), ...]

# Text β†’ brain regions
ckpt = hf_hub_download(REPO, "text2roi_combined_v4.pt")
roi_map = predict_text("I am terrified of the dark", ckpt)
print(top_rois(roi_map, n=5))
# β†’ [('Amygdala_L', 0.xx), ('AI', 0.xx), ('dACC', 0.xx), ...]

CLI

# Audio
python predict.py audio clip.wav --model text2roi_whisper_v4.pt --top 10

# Text
python predict.py text "watching a spider crawl toward me" --model text2roi_combined_v4.pt

# Text + Audio
python predict.py combined "narration text" clip.wav

Training Details

  • Corpus: 2.73M TRs Β· 289 GB Β· 4,480 sessions (CNeuroMod Friends, Narratives, LPP, HCP, language fMRI, Cowen-Keltner)
  • Brain space: fsaverage5 + subcortical, 28,444 vertices β†’ 56 ROI parcellation (HCP MMP1.0)
  • Architecture: Linear(in_dimβ†’1024) β†’ GELU β†’ Dropout β†’ LayerNorm β†’ Linear(1024β†’512) β†’ GELU β†’ Dropout β†’ Linear(512β†’56)
  • Loss: Pearson R (+ anchor ranking for qwen3 model)
  • Epochs: 120 per model
  • Eval metric: Mean Pearson R across 56 ROIs, identical to Algonauts 2025

Requirements

torch>=2.0
transformers>=4.40
huggingface_hub>=0.23
numpy>=1.24
librosa>=0.10   # for audio loading

Citation / Contact

Built by Amphora.

If you use this in research, please cite the HuggingFace repo URL.