File size: 10,604 Bytes
90480f5 ddedd58 90480f5 ddedd58 8209399 ddedd58 8209399 ddedd58 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 | ---
license: mit
library_name: pytorch
tags:
- computer-vision
- affective-computing
- neuroscience
- image-regression
- vgg16
- pavlovian-conditioning
- valence-prediction
---
# Visual-Valence Model (VCA)
> A deep neural network model of associative emotional (Pavlovian fear) learning.
Companion model repository for **["Associative Emotional Learning in Convolutional Neural Networks"](https://arxiv.org/abs/2607.19327)**
(Leem, Keil, Ding & Fang; *Neural Computation*, in press).
- 📄 Paper: https://arxiv.org/abs/2607.19327
- 💻 Code: https://github.com/lab-smile/FearConditioningAI
> **Note:** This model is a research artifact for computational/cognitive neuroscience, released to
> reproduce and extend the paper's findings. It is **not** a general-purpose sentiment/emotion API and
> is not validated for clinical, diagnostic, or production affective-computing use.
---
## Model description
The Visual-Valence Model predicts the affective **valence** of a visual scene (1 = extreme displeasure,
9 = extreme pleasure) and reproduces hallmarks of Pavlovian (fear) conditioning when a novel, initially
neutral stimulus is repeatedly paired with an emotionally-charged one.
The architecture (`Visual_Cortex_Amygdala` in [`models/VGG_Model.py`](https://github.com/lab-smile/FearConditioningAI/blob/main/models/VGG_Model.py))
has three components, loosely modeling the primate visual/affective pathway:
| Component | Role | Implementation |
|---|---|---|
| **Visual Cortex module** ("High Road") | Ventral visual stream | VGG-16 (Simonyan & Zisserman, 2015), ImageNet-pretrained, **frozen** |
| **Shortcut Pathway** ("Middle Road") | Fast, low-resolution route from early vision to affect circuitry | Multi-scale pooling of VGG-16's early layers (layer index 10), combined via an Efficient Channel Attention (ECA) module (Wang et al., 2020) |
| **Valence Module** | Combines both pathways into a scalar valence judgment | Two fully-connected layers (amygdala LA/CE-nuclei- and OFC-inspired) + one sigmoid output unit |
The High Road output (4096-d) and Middle Road output (512-d) are concatenated (4608-d) and passed through
the Valence Module's fully connected layers to a single sigmoid unit, which is linearly rescaled from
`[0, 1]` to the `[1, 9]` IAPS valence scale at evaluation time.
### Five checkpoints, one training pipeline
This repository hosts every checkpoint along the training pipeline described in the
[GitHub README](https://github.com/lab-smile/FearConditioningAI#training) — same architecture
(`Visual_Cortex_Amygdala`) throughout, so any of them can be loaded with the same code:
| File | Stage | Description |
|---|---|---|
| `vca_ckvideo_batch128_lr2e-5_epoch20.pth` | 0 | Trained from scratch on the Cowen & Keltner Videoframe dataset |
| `vca_IAPS_batch10_lr2e-4_epoch23.pth` | 1 | Fine-tuned on full-size IAPS images |
| `base_model_vca_IAPS_quadrant.pth` | 2 | Fine-tuned to the quadrant-cropped input layout — **pre-conditioning**: has learned to decode valence from natural scenes (the US) but has never seen the conditioned stimulus (CS, a Gabor patch) |
| `base_model_conditioned_orientation_epoch1.pth` | 3, epoch 1 | First epoch of Pavlovian conditioning (CS+ Gabor patch paired with pleasant/unpleasant IAPS US); early/under-trained, kept for provenance |
| `base_model_conditioned_orientation_epoch100.pth` | 3, epoch 100 | **Post-conditioning (final)** — used throughout the paper's conditioning/generalization/representational-alignment analyses |
Comparing the pre- (Stage 2) and post-conditioning (Stage 3, epoch 100) checkpoints' responses to the CS
alone is what reveals the learned CS→valence association (and, at the representation level, the
increasing CS/US alignment reported in the paper).
---
## Intended use
- Reproducing the paper's Pavlovian-conditioning experiments and figures.
- Extending the model to new conditioning paradigms, stimulus sets, or ablations (e.g. the
attention-free variant, `Visual_Cortex_Amygdala_wo_Attention`) for **research purposes**.
- Studying how a two-pathway (ventral-stream + shortcut) architecture with attention reproduces
behavioral/representational signatures of associative emotional learning.
- Feature extraction / representational analysis (see `Channel_Activity_Extraction.py`,
`Manifold_Visualization.py`, `SVM_Analysis_*.py` in the GitHub repo) for downstream neuroscience analyses.
**Out of scope:** general-purpose sentiment analysis, affect recognition on non-IAPS-like natural images,
clinical/diagnostic use, or any decision-making about real individuals' emotional states.
---
## Limitations
- **Frozen ImageNet backbone.** The High Road (VGG-16) is never fine-tuned, so the model inherits
ImageNet's visual biases; only the shortcut pathway and valence module are trained on affective data.
- **Narrow, licensed training data.** IAPS is a relatively small, curated stimulus set under a
data-use/confidentiality agreement (not redistributed with the code or this checkpoint); the model's
generalization to naturalistic, in-the-wild imagery is untested.
- **Two-alternative conditioning paradigm.** The conditioning stage pairs exactly two Gabor orientations
(45°, 135°) with unpleasant/pleasant IAPS images; behavior on other CS types has not been evaluated.
- **Not validated against individual human subjects.** Comparisons to human data in the paper are at the
group/aggregate level; the model is not a predictive model of any specific person's affective response.
- **Single scalar output.** The model predicts only valence (pleasant ↔ unpleasant), not arousal or
discrete emotion categories.
---
## Evaluation metrics
Model quality is reported using:
- **Pearson correlation (R / R²)** between predicted and ground-truth (SAM-rated) valence, computed by
`reg_eval_model` / `cond_eval_model` in [`utils.py`](https://github.com/lab-smile/FearConditioningAI/blob/main/utils.py).
- **Mean-squared error (MSE)** between predicted and ground-truth valence (rescaled to the 1–9 scale).
- Post-conditioning, generalization and CS/US **representational alignment** (single-unit tuning and
population-level RSA/t-SNE, via `Channel_Activity_Extraction.py`, `Manifold_Visualization.py`,
`SVM_Analysis_Emotion.py`, `SVM_Analysis_Before_After.py`) are used to assess whether conditioning
reproduces human associative-learning signatures.
The full quantitative results (per-stage R/MSE, generalization curves, and alignment statistics) are
reported in the paper's tables/figures — see https://arxiv.org/abs/2607.19327.
---
## Training dataset
Training proceeds in stages, each building on the previous stage's checkpoint (see the GitHub README's
[Training](https://github.com/lab-smile/FearConditioningAI#training) section for exact commands):
| Stage | Dataset | Purpose |
|---|---|---|
| 0 | Cowen & Keltner (2017) Videoframe dataset (2,185 emotion-eliciting video clips, one frame sampled every 10th frame) | Pretrain valence regression from scratch on natural scenes |
| 1 | International Affective Picture System (IAPS; Bradley & Lang, 1994/2007), full-size images, 8:1:1 train/val/test split | Adapt to the US stimuli used in conditioning |
| 2 | IAPS, quadrant-cropped layout | Fine-tune to the spatial layout used during conditioning |
| 3 | IAPS (US, in the 4th quadrant) × Gabor patches (CS, in the 2nd quadrant; 45°/135° orientation, varying spatial frequency/contrast, generated via `Gabor4Seowung.m`) | Pavlovian conditioning: pair CS+ with pleasant/unpleasant US |
Labels are human valence ratings on a 1–9 scale (Self-Assessment Manikin; Bradley & Lang, 1994).
**The IAPS images themselves are not redistributed** with the code or this model repository due to a
data-use/confidentiality agreement — obtain access to IAPS independently to reproduce training from
scratch. The Gabor-patch CS stimuli are procedurally generated and have no such restriction.
Note the input layout differs by stage: Stage 0/1 checkpoints expect a full-frame natural image
(resize + normalize only), while Stage 2/3 checkpoints expect the quadrant-cropped layout (see
[Data preprocessing](https://github.com/lab-smile/FearConditioningAI#data-preprocessing) in the GitHub
README) — feeding a full-frame image to a Stage 2/3 checkpoint (or vice versa) will not reproduce the
reported behavior.
---
## How to use
This is a plain PyTorch checkpoint (not a `transformers` model), so inference requires the model class
definition from the companion GitHub repository. See [`inference_example.py`](inference_example.py) in
this repository for a complete, runnable example. In short:
```bash
git clone https://github.com/lab-smile/FearConditioningAI.git
cd FearConditioningAI
pip install -r requirements.txt # or: conda env create -f environment-<platform>.yml
```
```python
import torch
from huggingface_hub import hf_hub_download
from models.VGG_Model import Visual_Cortex_Amygdala
repo_id = "smilelab/visual-valence-model"
# swap in any filename from the checkpoint table above, e.g. "vca_ckvideo_batch128_lr2e-5_epoch20.pth"
ckpt_path = hf_hub_download(repo_id=repo_id, filename="base_model_conditioned_orientation_epoch100.pth")
model = Visual_Cortex_Amygdala()
checkpoint = torch.load(ckpt_path, map_location="cpu", weights_only=False)
model.load_state_dict(checkpoint["state_dict"], strict=False)
model.eval()
```
See `inference_example.py` for image preprocessing (resize/normalize + quadrant placement of the CS/US)
and how to rescale the model's sigmoid output back to the 1–9 valence scale.
---
## Citation
If you use this model, please cite the paper:
```bibtex
@article{leem2026associative,
title = {Associative Emotional Learning in Convolutional Neural Networks},
author = {Leem, Seowung and Keil, Andreas and Ding, Mingzhou and Fang, Ruogu},
journal = {Neural Computation},
year = {2026},
note = {in press},
eprint = {2607.19327},
archivePrefix = {arXiv},
url = {https://arxiv.org/abs/2607.19327}
}
```
Please also cite the datasets and methods this model builds on (IAPS, Cowen & Keltner Videoframe, SAM,
VGG-16, ImageNet, ECA-Net, Rescorla-Wagner) — full references in the
[GitHub README's Citations section](https://github.com/lab-smile/FearConditioningAI#citations).
---
## License
This model is released under the **MIT License**, matching the
[GitHub repository](https://github.com/lab-smile/FearConditioningAI/blob/main/LICENSE).
## Contact
| Name | Email |
|---|---|
| Seowung Leem | leem.s@ufl.edu |
| Dr. Ruogu Fang | ruogu.fang@bme.ufl.edu |
|