Amphora NeuroText v4 β Audio & Text β Brain ROI Activation
Small MLP brain encoder: predict which brain regions activate in response to any audio or text β no fMRI required at inference time.
Trained exclusively on real naturalistic fMRI data.
This is a demo, not representative of the full model**.
Predicts activation across 56 canonical brain ROIs (HCP MMP1.0 parcellation).
Headline Result
Audio model (Whisper v4) beats TRIBE v2 β the Meta AI model that won the Algonauts 2025 competition β by +4.2% on a 23-subject cross-subject holdout:
| Model | Held-out R | Notes |
|---|---|---|
| NeuroText Whisper v4 | 0.257 | Single shared model, no per-subject fine-tuning |
| TRIBE v2 (Meta AI) | 0.215 | Published; per-subject fine-tuned; video+audio+text |
All 7 brain networks exceeded TRIBE v2 at a fraction of the compute. Biggest leads: Frontal (+0.209), Default Mode (+0.221), Subcortical (+0.234). These predictions are made on a coarser 56 ROI level, future work will be devoted towards increasing prediction fidelity.
Models
| File | Input | Val R | Holdout R | Notes |
|---|---|---|---|---|
text2roi_whisper_v4.pt |
Audio (Whisper-large-v3, 1280d) | 0.217 | 0.257 | Recommended for audio. Beats TRIBE v2 (+4.2%) |
text2roi_combined_v4.pt |
Text+Audio (3840d, modality dropout) | 0.192 | β | Text or combined inference; recommended for text |
text2roi_qwen3_v8.pt |
Text (Qwen3-Embedding-4B, 2560d) | 0.115 | β | Text-only; cross-dataset generalization improving |
Val R = mean Pearson R across 56 ROIs on held-out subjects (honest β see below). Holdout R = independent eval on 23 never-seen subjects.
Note on previous models (v2/v3): Models in this repo before July 2026 reported inflated val_R values (0.239β0.413) due to within-subject train/val splits. Those numbers are not comparable to v4. The v4 models use honest per-subject z-scoring and per-subject holdout splits.
Why These Numbers Are Honest
Previous NeuroText versions had inflated val_R from a within-subject split: the model saw the same subjects in both train and val, and learned subject-level baseline BOLD activations. Cross-subject, that memorized baseline is useless.
v4 fixes both issues:
- Per-subject z-scoring β each subject's ROI activations are z-scored independently before training, so the model learns stimulus-driven fluctuations rather than who each subject is.
- Per-subject holdout split β 15% of subjects per dataset are excluded from training entirely. Zero subject overlap guaranteed.
- Single shared model β no per-subject adaptation at inference. TRIBE requires fine-tuning on each test subject; Amphora does not.
Audio Model β Top ROIs (held-out eval)
| ROI | R | vs TRIBE |
|---|---|---|
| ACC | 0.438 | +0.368 |
| STG | 0.422 | +0.302 |
| Thalamus | 0.408 | +0.358 |
| V1 | 0.393 | +0.073 |
| LP_R | 0.380 | +0.290 |
| mPFC_dmn | 0.379 | +0.299 |
| AI | 0.343 | +0.253 |
| HPC_L | 0.337 | +0.247 |
| V2 | 0.336 | +0.036 |
| dACC | 0.330 | +0.250 |
53/56 ROIs R>0.10 Β· 41/56 R>0.20 Β· 20/56 R>0.30
Quick Start
from huggingface_hub import hf_hub_download
from predict import predict_audio, predict_text, top_rois
REPO = "ffh92r32rm0/Amphora_NeuroText"
# Audio β brain regions (recommended)
ckpt = hf_hub_download(REPO, "text2roi_whisper_v4.pt")
roi_map = predict_audio("clip.wav", ckpt)
print(top_rois(roi_map, n=5))
# β [('ACC', 0.44), ('STG', 0.42), ('Thalamus', 0.41), ...]
# Text β brain regions
ckpt = hf_hub_download(REPO, "text2roi_combined_v4.pt")
roi_map = predict_text("I am terrified of the dark", ckpt)
print(top_rois(roi_map, n=5))
# β [('Amygdala_L', 0.xx), ('AI', 0.xx), ('dACC', 0.xx), ...]
CLI
# Audio
python predict.py audio clip.wav --model text2roi_whisper_v4.pt --top 10
# Text
python predict.py text "watching a spider crawl toward me" --model text2roi_combined_v4.pt
# Text + Audio
python predict.py combined "narration text" clip.wav
Training Details
- Corpus: 2.73M TRs Β· 289 GB Β· 4,480 sessions (CNeuroMod Friends, Narratives, LPP, HCP, language fMRI, Cowen-Keltner)
- Brain space: fsaverage5 + subcortical, 28,444 vertices β 56 ROI parcellation (HCP MMP1.0)
- Architecture: Linear(in_dimβ1024) β GELU β Dropout β LayerNorm β Linear(1024β512) β GELU β Dropout β Linear(512β56)
- Loss: Pearson R (+ anchor ranking for qwen3 model)
- Epochs: 120 per model
- Eval metric: Mean Pearson R across 56 ROIs, identical to Algonauts 2025
Requirements
torch>=2.0
transformers>=4.40
huggingface_hub>=0.23
numpy>=1.24
librosa>=0.10 # for audio loading
Citation / Contact
Built by Amphora.
If you use this in research, please cite the HuggingFace repo URL.
- Downloads last month
- 65