Amphora NeuroText v4 β€” Audio & Text β†’ Brain ROI Activation

Small MLP brain encoder: predict which brain regions activate in response to any audio or text β€” no fMRI required at inference time.

Open In Colab

Trained exclusively on real naturalistic fMRI data.

This is a demo, not representative of the full model**.

Predicts activation across 56 canonical brain ROIs (HCP MMP1.0 parcellation).


Headline Result

Audio model (Whisper v4) beats TRIBE v2 β€” the Meta AI model that won the Algonauts 2025 competition β€” by +4.2% on a 23-subject cross-subject holdout:

Model Held-out R Notes
NeuroText Whisper v4 0.257 Single shared model, no per-subject fine-tuning
TRIBE v2 (Meta AI) 0.215 Published; per-subject fine-tuned; video+audio+text

All 7 brain networks exceeded TRIBE v2 at a fraction of the compute. Biggest leads: Frontal (+0.209), Default Mode (+0.221), Subcortical (+0.234). These predictions are made on a coarser 56 ROI level, future work will be devoted towards increasing prediction fidelity.


Models

File Input Val R Holdout R Notes
text2roi_whisper_v4.pt Audio (Whisper-large-v3, 1280d) 0.217 0.257 Recommended for audio. Beats TRIBE v2 (+4.2%)
text2roi_combined_v4.pt Text+Audio (3840d, modality dropout) 0.192 β€” Text or combined inference; recommended for text
text2roi_qwen3_v8.pt Text (Qwen3-Embedding-4B, 2560d) 0.115 β€” Text-only; cross-dataset generalization improving

Val R = mean Pearson R across 56 ROIs on held-out subjects (honest β€” see below). Holdout R = independent eval on 23 never-seen subjects.

Note on previous models (v2/v3): Models in this repo before July 2026 reported inflated val_R values (0.239–0.413) due to within-subject train/val splits. Those numbers are not comparable to v4. The v4 models use honest per-subject z-scoring and per-subject holdout splits.


Why These Numbers Are Honest

Previous NeuroText versions had inflated val_R from a within-subject split: the model saw the same subjects in both train and val, and learned subject-level baseline BOLD activations. Cross-subject, that memorized baseline is useless.

v4 fixes both issues:

  1. Per-subject z-scoring β€” each subject's ROI activations are z-scored independently before training, so the model learns stimulus-driven fluctuations rather than who each subject is.
  2. Per-subject holdout split β€” 15% of subjects per dataset are excluded from training entirely. Zero subject overlap guaranteed.
  3. Single shared model β€” no per-subject adaptation at inference. TRIBE requires fine-tuning on each test subject; Amphora does not.

Audio Model β€” Top ROIs (held-out eval)

ROI R vs TRIBE
ACC 0.438 +0.368
STG 0.422 +0.302
Thalamus 0.408 +0.358
V1 0.393 +0.073
LP_R 0.380 +0.290
mPFC_dmn 0.379 +0.299
AI 0.343 +0.253
HPC_L 0.337 +0.247
V2 0.336 +0.036
dACC 0.330 +0.250

53/56 ROIs R>0.10 Β· 41/56 R>0.20 Β· 20/56 R>0.30


Quick Start

from huggingface_hub import hf_hub_download
from predict import predict_audio, predict_text, top_rois

REPO = "ffh92r32rm0/Amphora_NeuroText"

# Audio β†’ brain regions (recommended)
ckpt = hf_hub_download(REPO, "text2roi_whisper_v4.pt")
roi_map = predict_audio("clip.wav", ckpt)
print(top_rois(roi_map, n=5))
# β†’ [('ACC', 0.44), ('STG', 0.42), ('Thalamus', 0.41), ...]

# Text β†’ brain regions
ckpt = hf_hub_download(REPO, "text2roi_combined_v4.pt")
roi_map = predict_text("I am terrified of the dark", ckpt)
print(top_rois(roi_map, n=5))
# β†’ [('Amygdala_L', 0.xx), ('AI', 0.xx), ('dACC', 0.xx), ...]

CLI

# Audio
python predict.py audio clip.wav --model text2roi_whisper_v4.pt --top 10

# Text
python predict.py text "watching a spider crawl toward me" --model text2roi_combined_v4.pt

# Text + Audio
python predict.py combined "narration text" clip.wav

Training Details

  • Corpus: 2.73M TRs Β· 289 GB Β· 4,480 sessions (CNeuroMod Friends, Narratives, LPP, HCP, language fMRI, Cowen-Keltner)
  • Brain space: fsaverage5 + subcortical, 28,444 vertices β†’ 56 ROI parcellation (HCP MMP1.0)
  • Architecture: Linear(in_dimβ†’1024) β†’ GELU β†’ Dropout β†’ LayerNorm β†’ Linear(1024β†’512) β†’ GELU β†’ Dropout β†’ Linear(512β†’56)
  • Loss: Pearson R (+ anchor ranking for qwen3 model)
  • Epochs: 120 per model
  • Eval metric: Mean Pearson R across 56 ROIs, identical to Algonauts 2025

Requirements

torch>=2.0
transformers>=4.40
huggingface_hub>=0.23
numpy>=1.24
librosa>=0.10   # for audio loading

Citation / Contact

Built by Amphora.

If you use this in research, please cite the HuggingFace repo URL.

Downloads last month
65
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support