File size: 10,604 Bytes
90480f5
 
ddedd58
 
 
 
 
 
 
 
 
90480f5
ddedd58
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8209399
 
 
 
 
 
 
ddedd58
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
8209399
ddedd58
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
---
license: mit
library_name: pytorch
tags:
  - computer-vision
  - affective-computing
  - neuroscience
  - image-regression
  - vgg16
  - pavlovian-conditioning
  - valence-prediction
---

# Visual-Valence Model (VCA)

> A deep neural network model of associative emotional (Pavlovian fear) learning.

Companion model repository for **["Associative Emotional Learning in Convolutional Neural Networks"](https://arxiv.org/abs/2607.19327)**
(Leem, Keil, Ding & Fang; *Neural Computation*, in press).

- 📄 Paper: https://arxiv.org/abs/2607.19327
- 💻 Code: https://github.com/lab-smile/FearConditioningAI

> **Note:** This model is a research artifact for computational/cognitive neuroscience, released to
> reproduce and extend the paper's findings. It is **not** a general-purpose sentiment/emotion API and
> is not validated for clinical, diagnostic, or production affective-computing use.

---

## Model description

The Visual-Valence Model predicts the affective **valence** of a visual scene (1 = extreme displeasure,
9 = extreme pleasure) and reproduces hallmarks of Pavlovian (fear) conditioning when a novel, initially
neutral stimulus is repeatedly paired with an emotionally-charged one.

The architecture (`Visual_Cortex_Amygdala` in [`models/VGG_Model.py`](https://github.com/lab-smile/FearConditioningAI/blob/main/models/VGG_Model.py))
has three components, loosely modeling the primate visual/affective pathway:

| Component | Role | Implementation |
|---|---|---|
| **Visual Cortex module** ("High Road") | Ventral visual stream | VGG-16 (Simonyan & Zisserman, 2015), ImageNet-pretrained, **frozen** |
| **Shortcut Pathway** ("Middle Road") | Fast, low-resolution route from early vision to affect circuitry | Multi-scale pooling of VGG-16's early layers (layer index 10), combined via an Efficient Channel Attention (ECA) module (Wang et al., 2020) |
| **Valence Module** | Combines both pathways into a scalar valence judgment | Two fully-connected layers (amygdala LA/CE-nuclei- and OFC-inspired) + one sigmoid output unit |

The High Road output (4096-d) and Middle Road output (512-d) are concatenated (4608-d) and passed through
the Valence Module's fully connected layers to a single sigmoid unit, which is linearly rescaled from
`[0, 1]` to the `[1, 9]` IAPS valence scale at evaluation time.

### Five checkpoints, one training pipeline

This repository hosts every checkpoint along the training pipeline described in the
[GitHub README](https://github.com/lab-smile/FearConditioningAI#training) — same architecture
(`Visual_Cortex_Amygdala`) throughout, so any of them can be loaded with the same code:

| File | Stage | Description |
|---|---|---|
| `vca_ckvideo_batch128_lr2e-5_epoch20.pth` | 0 | Trained from scratch on the Cowen & Keltner Videoframe dataset |
| `vca_IAPS_batch10_lr2e-4_epoch23.pth` | 1 | Fine-tuned on full-size IAPS images |
| `base_model_vca_IAPS_quadrant.pth` | 2 | Fine-tuned to the quadrant-cropped input layout — **pre-conditioning**: has learned to decode valence from natural scenes (the US) but has never seen the conditioned stimulus (CS, a Gabor patch) |
| `base_model_conditioned_orientation_epoch1.pth` | 3, epoch 1 | First epoch of Pavlovian conditioning (CS+ Gabor patch paired with pleasant/unpleasant IAPS US); early/under-trained, kept for provenance |
| `base_model_conditioned_orientation_epoch100.pth` | 3, epoch 100 | **Post-conditioning (final)** — used throughout the paper's conditioning/generalization/representational-alignment analyses |

Comparing the pre- (Stage 2) and post-conditioning (Stage 3, epoch 100) checkpoints' responses to the CS
alone is what reveals the learned CS→valence association (and, at the representation level, the
increasing CS/US alignment reported in the paper).

---

## Intended use

- Reproducing the paper's Pavlovian-conditioning experiments and figures.
- Extending the model to new conditioning paradigms, stimulus sets, or ablations (e.g. the
  attention-free variant, `Visual_Cortex_Amygdala_wo_Attention`) for **research purposes**.
- Studying how a two-pathway (ventral-stream + shortcut) architecture with attention reproduces
  behavioral/representational signatures of associative emotional learning.
- Feature extraction / representational analysis (see `Channel_Activity_Extraction.py`,
  `Manifold_Visualization.py`, `SVM_Analysis_*.py` in the GitHub repo) for downstream neuroscience analyses.

**Out of scope:** general-purpose sentiment analysis, affect recognition on non-IAPS-like natural images,
clinical/diagnostic use, or any decision-making about real individuals' emotional states.

---

## Limitations

- **Frozen ImageNet backbone.** The High Road (VGG-16) is never fine-tuned, so the model inherits
  ImageNet's visual biases; only the shortcut pathway and valence module are trained on affective data.
- **Narrow, licensed training data.** IAPS is a relatively small, curated stimulus set under a
  data-use/confidentiality agreement (not redistributed with the code or this checkpoint); the model's
  generalization to naturalistic, in-the-wild imagery is untested.
- **Two-alternative conditioning paradigm.** The conditioning stage pairs exactly two Gabor orientations
  (45°, 135°) with unpleasant/pleasant IAPS images; behavior on other CS types has not been evaluated.
- **Not validated against individual human subjects.** Comparisons to human data in the paper are at the
  group/aggregate level; the model is not a predictive model of any specific person's affective response.
- **Single scalar output.** The model predicts only valence (pleasant ↔ unpleasant), not arousal or
  discrete emotion categories.

---

## Evaluation metrics

Model quality is reported using:

- **Pearson correlation (R / R²)** between predicted and ground-truth (SAM-rated) valence, computed by
  `reg_eval_model` / `cond_eval_model` in [`utils.py`](https://github.com/lab-smile/FearConditioningAI/blob/main/utils.py).
- **Mean-squared error (MSE)** between predicted and ground-truth valence (rescaled to the 1–9 scale).
- Post-conditioning, generalization and CS/US **representational alignment** (single-unit tuning and
  population-level RSA/t-SNE, via `Channel_Activity_Extraction.py`, `Manifold_Visualization.py`,
  `SVM_Analysis_Emotion.py`, `SVM_Analysis_Before_After.py`) are used to assess whether conditioning
  reproduces human associative-learning signatures.

The full quantitative results (per-stage R/MSE, generalization curves, and alignment statistics) are
reported in the paper's tables/figures — see https://arxiv.org/abs/2607.19327.

---

## Training dataset

Training proceeds in stages, each building on the previous stage's checkpoint (see the GitHub README's
[Training](https://github.com/lab-smile/FearConditioningAI#training) section for exact commands):

| Stage | Dataset | Purpose |
|---|---|---|
| 0 | Cowen & Keltner (2017) Videoframe dataset (2,185 emotion-eliciting video clips, one frame sampled every 10th frame) | Pretrain valence regression from scratch on natural scenes |
| 1 | International Affective Picture System (IAPS; Bradley & Lang, 1994/2007), full-size images, 8:1:1 train/val/test split | Adapt to the US stimuli used in conditioning |
| 2 | IAPS, quadrant-cropped layout | Fine-tune to the spatial layout used during conditioning |
| 3 | IAPS (US, in the 4th quadrant) × Gabor patches (CS, in the 2nd quadrant; 45°/135° orientation, varying spatial frequency/contrast, generated via `Gabor4Seowung.m`) | Pavlovian conditioning: pair CS+ with pleasant/unpleasant US |

Labels are human valence ratings on a 1–9 scale (Self-Assessment Manikin; Bradley & Lang, 1994).
**The IAPS images themselves are not redistributed** with the code or this model repository due to a
data-use/confidentiality agreement — obtain access to IAPS independently to reproduce training from
scratch. The Gabor-patch CS stimuli are procedurally generated and have no such restriction.

Note the input layout differs by stage: Stage 0/1 checkpoints expect a full-frame natural image
(resize + normalize only), while Stage 2/3 checkpoints expect the quadrant-cropped layout (see
[Data preprocessing](https://github.com/lab-smile/FearConditioningAI#data-preprocessing) in the GitHub
README) — feeding a full-frame image to a Stage 2/3 checkpoint (or vice versa) will not reproduce the
reported behavior.

---

## How to use

This is a plain PyTorch checkpoint (not a `transformers` model), so inference requires the model class
definition from the companion GitHub repository. See [`inference_example.py`](inference_example.py) in
this repository for a complete, runnable example. In short:

```bash
git clone https://github.com/lab-smile/FearConditioningAI.git
cd FearConditioningAI
pip install -r requirements.txt   # or: conda env create -f environment-<platform>.yml
```

```python
import torch
from huggingface_hub import hf_hub_download
from models.VGG_Model import Visual_Cortex_Amygdala

repo_id = "smilelab/visual-valence-model"
# swap in any filename from the checkpoint table above, e.g. "vca_ckvideo_batch128_lr2e-5_epoch20.pth"
ckpt_path = hf_hub_download(repo_id=repo_id, filename="base_model_conditioned_orientation_epoch100.pth")

model = Visual_Cortex_Amygdala()
checkpoint = torch.load(ckpt_path, map_location="cpu", weights_only=False)
model.load_state_dict(checkpoint["state_dict"], strict=False)
model.eval()
```

See `inference_example.py` for image preprocessing (resize/normalize + quadrant placement of the CS/US)
and how to rescale the model's sigmoid output back to the 1–9 valence scale.

---

## Citation

If you use this model, please cite the paper:

```bibtex
@article{leem2026associative,
  title   = {Associative Emotional Learning in Convolutional Neural Networks},
  author  = {Leem, Seowung and Keil, Andreas and Ding, Mingzhou and Fang, Ruogu},
  journal = {Neural Computation},
  year    = {2026},
  note    = {in press},
  eprint  = {2607.19327},
  archivePrefix = {arXiv},
  url     = {https://arxiv.org/abs/2607.19327}
}
```

Please also cite the datasets and methods this model builds on (IAPS, Cowen & Keltner Videoframe, SAM,
VGG-16, ImageNet, ECA-Net, Rescorla-Wagner) — full references in the
[GitHub README's Citations section](https://github.com/lab-smile/FearConditioningAI#citations).

---

## License

This model is released under the **MIT License**, matching the
[GitHub repository](https://github.com/lab-smile/FearConditioningAI/blob/main/LICENSE).

## Contact

| Name | Email |
|---|---|
| Seowung Leem | leem.s@ufl.edu |
| Dr. Ruogu Fang | ruogu.fang@bme.ufl.edu |