File size: 18,757 Bytes
312d325 6dbbe28 312d325 6dbbe28 bb02bb9 6dbbe28 bb02bb9 6dbbe28 9b20c58 6dbbe28 dca8f60 ab80480 dca8f60 6dbbe28 3ac289c 6dbbe28 cfcb7d9 6dbbe28 18a0bc6 6dbbe28 18a0bc6 6dbbe28 d9dab2f 6dbbe28 18a0bc6 6dbbe28 18a0bc6 385646d 6dbbe28 da2c028 6dbbe28 0a49d34 6dbbe28 89c5e0c 6dbbe28 89c5e0c 6dbbe28 f4bbc8a 6dbbe28 4a8d313 6dbbe28 f053c66 6dbbe28 8e04562 6dbbe28 46ba62f 6dbbe28 ac65ef9 6dbbe28 ac65ef9 6dbbe28 ac65ef9 da2c028 ac65ef9 6dbbe28 ac65ef9 aaf8182 6dbbe28 1d3a8dd da2c028 1d3a8dd 6dbbe28 da2c028 6dbbe28 da2c028 6dbbe28 25ab075 6dbbe28 9b20c58 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 284 285 286 287 288 289 290 291 292 293 294 295 296 297 298 299 300 301 302 303 304 305 306 307 308 309 310 311 312 313 314 315 316 317 318 319 320 321 322 323 324 325 326 327 328 329 330 331 332 333 334 335 336 337 338 339 340 341 342 343 344 345 346 347 348 349 350 351 352 353 354 355 356 357 358 359 360 361 362 363 364 365 366 367 368 369 370 371 372 373 374 375 376 377 378 379 380 381 382 383 384 385 386 387 388 389 390 391 392 393 394 395 396 397 398 399 400 401 402 403 404 405 406 407 408 | ---
license: apache-2.0
tags:
- pytorch
- computer-vision
- self-supervised-learning
- simclr
- resnet18
- egocentric-vision
- eccentricity
- visual-neuroscience
- vedb
- arxiv:2607.19316
---
# VEDB SimCLR ResNet-18 β Baseline
This repository contains the **Baseline SimCLR ResNet-18 checkpoint** from:
**Diaz, D. M., & Henderson, M. M. (2026). _Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field._ Proceedings of the Conference on Cognitive Computational Neuroscience 2026.**
**DOI:** `10.32470/0416gfsq`<br>
**arXiv:** `2607.19316`<br>
**Contributed Talk:** [CCN 2026 presentation on YouTube](https://www.youtube.com/watch?v=Lb4S3FWqd2M&t=2545s)
This model is part of the **Eccentricity-Constrained SimCLR Models (VEDB)** [collection](https://hf.co/collections/DM-Diaz/eccentricity-constrained-simclr-models-vedb), containing checkpoints pretrained under four visual-field conditions: **Baseline, Fovea-Gaze, Periph, and Periph-NF**.
## Model Description
This model uses a **ResNet-18 visual encoder pretrained with SimCLR self-supervised contrastive learning** on naturalistic egocentric imagery with synchronized human gaze data from the [**Visual Experience Dataset (VEDB)**](https://jov.arvojournals.org/article.aspx?articleid=2802101).
The associated study investigated whether constraining visual experience to different portions of the visual field produces systematic differences in learned representations, downstream task performance, and alignment with human visual cortex.
### Baseline Condition
<p align="center">
<img
src="https://huggingface.co/DM-Diaz/VEDB-SimCLR-ResNet18-Baseline/resolve/main/VEDB_Frame_Manipulation.png"
alt="Example VEDB frames under the Baseline, Fovea-Gaze, Periph, and Periph-NF training conditions"
width="850"
>
</p>
The **Baseline** condition serves as the full-field reference model. It was pretrained on the original, unedited `224 Γ 224` VEDB frames following the common frame preprocessing pipeline and therefore received **no additional central or peripheral visual-field restriction**.
The associated model variants manipulate the same source frames to isolate different forms of visual-field information:
- **Fovea-Gaze:** gaze-centered central-only input
- **Periph:** peripheral-only input produced by masking the gaze-centered central region
- **Periph-NF:** peripheral-only input with a [NeuroFovea](https://github.com/ArturoDeza/NeuroFovea) transform applied before central masking
## Release Status
| Component | Status |
| --- | --- |
| Pretrained checkpoint | Available |
| Model card | Available |
| Training code | Forthcoming |
| Evaluation code | Forthcoming |
| VEDB imagery | Not redistributed; available via [Databrary](https://www.databrary.org/volume/1612) |
The complete training and evaluation codebase is currently being consolidated and documented and will be linked here upon public release.
The checkpoint is being released in advance of the codebase to provide access to the model artifact used in the published study.
## Technical Provenance Note and Discrepancies
Technical provenance note: The checkpoint metadata, architecture, training parameters, and implementation details documented in this model card have been re-verified against the released model checkpoint and, where available, the original training code and launch configuration. If a technical detail concerning the released model artifact differs between the associated paper and this model card, the model card should be treated as the authoritative description of the released checkpoint and its implementation. The associated paper remains the primary source for the study's scientific analyses, results, and interpretation.
## Checkpoint
**File:**
`simclr_resnet18_baseline_epoch120.pth.tar`
This repository provides the original PyTorch training checkpoint from
epoch 120 of Baseline SimCLR pretraining.
The checkpoint is serialized as a dictionary containing:
| Key | Contents |
| --- | --- |
| `epoch` | Final training epoch (`120`) |
| `arch` | Backbone architecture (`resnet18`) |
| `state_dict` | Model parameters and registered buffers |
| `optimizer` | Adam optimizer state at the time of saving |
The `state_dict` contains **124 entries** and includes both the ResNet-18
encoder and SimCLR projection head.
Model parameters use the `backbone.*` namespace. The projection head is
stored as:
- `backbone.fc.0`: `Linear(512, 512)`
- `backbone.fc.2`: `Linear(512, 128)`
The checkpoint therefore contains the full SimCLR model state rather than
encoder weights alone.
## Architecture
<p align="center">
<img src="./architecture_overview.png" alt="Architecture overview" width="100%">
</p>
<p align="center">
<em>Overview of the VEDB preprocessing, SimCLR pretraining, downstream linear probes, and voxelwise encoding workflow.</em>
</p>
| Component | Specification |
| --- | --- |
| Backbone | ResNet-18 |
| Framework | PyTorch |
| Learning paradigm | Self-supervised contrastive learning |
| Objective | SimCLR / NT-Xent |
| Input resolution | `224 Γ 224` |
| Backbone representation | 512-dimensional |
| Projection head | `Linear(512, 512) β ReLU β Linear(512, 128)` |
| Projection dimension | 128 |
The standard ResNet-18 classification layer was replaced during SimCLR
pretraining by a two-layer projection head:
```python
nn.Sequential(
nn.Linear(512, 512),
nn.ReLU(),
nn.Linear(512, 128),
)
```
## Training Data
### Visual Experience Dataset (VEDB)
The model was pretrained using imagery from the **Visual Experience Dataset (VEDB)**, a large-scale dataset of naturalistic egocentric experience containing more than 200 hours of integrated egocentric video, eye-movement, and odometry recordings.
**VEDB resources:**
- **Dataset:** [VEDB on Databrary](https://www.databrary.org/volume/1612)
- **Dataset paper:** [Greene et al. (2024), _The Visual Experience Dataset: Over 200 recorded hours of integrated eye movement, odometry, and egocentric video_](https://jov.arvojournals.org/article.aspx?articleid=2802101)
- **Project resources:** [VEDB on OSF](https://osf.io/2gdkb/overview)
The original VEDB imagery is **not redistributed through this repository**. Researchers wishing to reproduce training should obtain VEDB through the official distribution and comply with its applicable access and usage requirements.
### Dataset Construction
A total of **717 VEDB sessions** were initially retrieved through Databrary. Sessions without synchronized gaze data were excluded, leaving **514 sessions** for processing and analysis.
Within task-relevant portions of each retained session:
- Frames were sampled every **2 seconds**.
- Sampling used the native **25 FPS** video rate.
- A maximum of **1,000 frames per session** was sampled.
- The resulting SimCLR dataset contained **433,564 frames**.
All data splits were performed at the **video-session level** to prevent leakage from temporally adjacent and environmentally correlated frames.
### SimCLR Dataset Split
| Split | Sessions | Frames | Frame proportion |
| --- | ---: | ---: | ---: |
| Train | 455 | 377,462 | 87.06% |
| Validation | 28 | 26,026 | 6.00% |
| Test | 31 | 30,076 | 6.94% |
| **Total** | **514** | **433,564** | **100%** |
The split was constructed as an approximately **80/10/10 session-level split** using stratification to balance task-label representation. Because sessions contain different numbers of sampled frames, the resulting frame percentages differ from the session-level proportions.
Validation and test sessions were held out from SimCLR representation learning.
## Frame Preprocessing
Sampled VEDB frames were processed using a deterministic common pipeline:
1. Decode the sampled video frame.
2. Convert the image to RGB.
3. Bicubic resize to **256 px**.
4. Center crop to **`224 Γ 224`**.
For the **Baseline condition**, no eccentricity-specific transformation was applied after this preprocessing.
Condition-specific image construction for Fovea-Gaze, Periph, and Periph-NF occurred before the shared SimCLR augmentation pipeline.
## SimCLR Pretraining
The four VEDB conditions used the same SimCLR architecture ([PyTorch-SimCLR](https://github.com/sthalles/SimCLR)), optimization procedure, and augmentation pipeline. They differed only in the visual-field manipulation applied to the source imagery before SimCLR augmentation.
| Hyperparameter | Value |
| --- | --- |
| Backbone | ResNet-18 |
| Input size | `224 Γ 224` |
| Epochs | 120 |
| Batch size | 512 |
| Optimizer | Adam |
| Learning rate | `6 Γ 10^-4` |
| Weight decay | `1 Γ 10^-4` |
| Loss | NT-Xent |
| Temperature (Ο) | `0.07` |
| Learning-rate schedule | CosineAnnealingLR (scheduler stepping begins after epoch 10) |
| Projection head | `Linear(512,512) β ReLU β Linear(512,128)` |
| Mixed precision | FP16 |
### SimCLR Augmentations
The common SimCLR augmentation pipeline included:
- random resized cropping,
- random horizontal flipping,
- color jitter (`p = 0.8`),
- grayscale conversion (`p = 0.2`), and
- Gaussian blur with `Ο ~ U(0.1, 2.0)`.
The same augmentation pipeline was used across all four VEDB conditions and was applied **after condition-specific frame construction**. See [PyTorch-SimCLR](https://github.com/sthalles/SimCLR) for further SimCLR implementation details.
## Evaluation
Following SimCLR pretraining, the frozen ResNet-18 backbone was evaluated using linear probes for in-domain and out-of-domain classification and voxelwise encoding models for neural prediction.
### Comparative Evaluation Results
The table below reproduces the summary metrics reported in the associated paper across all VEDB-trained conditions and reference models. **Rows corresponding to this repository's Baseline checkpoint are bolded.**
| Task | Condition | Val Loss | Top-1 (%) | Top-5 (%) | Best Macro-F1 (%) |
| --- | --- | ---: | ---: | ---: | ---: |
| SimCLR | **Baseline** | **0.4331** | **87.60** | β | β |
| SimCLR | Fovea-Gaze | 0.3749 | 90.43 | β | β |
| SimCLR | Periph-NF | 0.4548 | 90.04 | β | β |
| SimCLR | Periph | 0.4545 | 89.26 | β | β |
| In-Domain | **Baseline** | **0.9811** | β | β | **42.17** |
| In-Domain | Fovea-Gaze | 1.2031 | β | β | 43.64 |
| In-Domain | Periph-NF | 1.3090 | β | β | 30.93 |
| In-Domain | Periph | 1.0623 | β | β | 36.56 |
| In-Domain | STL-10 | 1.6666 | β | β | 25.41 |
| In-Domain | ImageNet-100 | 1.2342 | β | β | 41.23 |
| In-Domain | ImageNet-1K | 0.9713 | β | β | 43.33 |
| VGGFace2 | **Baseline** | **7.8101** | **5.21** | **11.73** | **3.26** |
| VGGFace2 | Fovea-Gaze | 7.9104 | 4.58 | 10.76 | 2.70 |
| VGGFace2 | Periph-NF | 8.0232 | 3.39 | 8.17 | 1.90 |
| VGGFace2 | Periph | 8.1681 | 2.54 | 6.39 | 1.35 |
| VGGFace2 | STL-10 | 6.9973 | 9.55 | 18.96 | 7.43 |
| VGGFace2 | ImageNet-100 | 6.7985 | 10.77 | 21.07 | 8.71 |
| VGGFace2 | ImageNet-1K | 6.7964 | 10.74 | 21.08 | 8.77 |
| Places365 | **Baseline** | **3.9690** | **25.63** | **51.90** | **23.16** |
| Places365 | Fovea-Gaze | 4.2347 | 21.86 | 46.21 | 19.14 |
| Places365 | Periph-NF | 4.2621 | 20.51 | 44.58 | 17.86 |
| Places365 | Periph | 4.2671 | 20.26 | 44.10 | 17.65 |
| Places365 | STL-10 | 3.8281 | 26.57 | 53.47 | 24.82 |
| Places365 | ImageNet-100 | 3.9207 | 24.99 | 51.21 | 23.32 |
| Places365 | ImageNet-1K | 3.6264 | 30.17 | 58.46 | 28.36 |
**Note:** SimCLR Top-1 is computed from the self-supervised contrastive objective and is not directly comparable to downstream supervised classification accuracy. For downstream tasks, the pretrained ResNet-18 backbone was **frozen** and only a linear classifier was trained; the backbone weights were **not fine-tuned**. Classifier checkpoints were selected by best validation Macro-F1. In-domain Top-1 accuracy is omitted because label imbalance across frames can make accuracy misleading; Macro-F1 is reported as the primary class-balanced metric. STL-10, ImageNet-100, and ImageNet-1K are treated as out-of-domain baselines because they were not pretrained on VEDB.
For in-domain classification, Macro-F1 was used as the primary class-balanced metric because of label imbalance across VEDB frame categories.
## Neural Encoding Evaluation
The pretrained model was additionally evaluated using voxelwise encoding models of human fMRI responses from the **Natural Scenes Dataset (NSD)**.
NSD contains 7T whole-brain fMRI responses to complex natural scenes. The analysis used data from **8 human participants**.
For each model:
- features were extracted from `Conv1`, `Layer1.1`, `Layer2.1`, `Layer3.1`, `Layer4.1`, and `Avgpool`,
- convolutional features were spatially downsampled,
- PCA was used to retain the top 200 components per feature set,
- features were concatenated and z-scored across images, and
- regularized L2 linear regression was used to predict individual voxel responses.
For each participant, the **1,000 NSD images shared across all participants** served as the held-out test set, while the remaining **9,000 images** viewed by that participant were used to fit the encoding models.
Importantly, **original intact NSD images were presented to every pretrained model during encoding evaluation**. The Baseline, Fovea-Gaze, Periph, and Periph-NF visual-field transformations were applied during SimCLR pretraining and were **not reapplied to NSD stimuli at the encoding stage**.
Encoding performance was quantified using held-out voxelwise `RΒ²`.
For complete ROI-level prediction accuracy, statistical comparisons, variance-partitioning analyses, and comparisons across eccentricity conditions, see the associated paper.
## Intended Use
This checkpoint is provided primarily for research involving:
- self-supervised visual representation learning,
- egocentric visual experience,
- central versus peripheral information processing,
- visual-field eccentricity,
- transfer learning and linear probing,
- computational modeling of visual cortex, and
- model-to-brain comparisons.
The checkpoint may also be used as a pretrained ResNet-18 initialization for methodological extensions or comparisons with alternative visual-field manipulations.
## Out-of-Scope Use
This model was developed as a **research representation-learning model** and was not designed or validated as:
- a production image-classification system,
- a general-purpose computer-vision foundation model,
- a biological simulation of the human visual system, or
- a system for making decisions about individuals.
The Baseline model contains no explicit simulation of retinal or cortical eccentricity.
## Limitations
VEDB consists of naturalistic first-person visual experience and is consequently more temporally correlated and semantically constrained than large curated computer-vision datasets.
Only **one SimCLR pretraining run per VEDB condition** was used in the published study. These checkpoints therefore do not characterize variation across independent pretraining seeds.
The learned representations are specific to the architecture, training objective, augmentations, data-sampling procedure, visual-field manipulation, and preprocessing choices used in the study. Alternative implementations may produce different representations or downstream performance.
These weights should therefore be interpreted as reproducible artifacts of the published experimental conditions. The authors do not make the claim that the particular implementation is the uniquely optimal method for modeling visual-field eccentricity.
## Loading the Model
The checkpoint contains the complete SimCLR model state, including the
ResNet-18 encoder and projection head.
```python
import torch
import torch.nn as nn
from torchvision.models import resnet18
class SimCLRResNet18(nn.Module):
def __init__(self):
super().__init__()
self.backbone = resnet18(weights=None)
self.backbone.fc = nn.Sequential(
nn.Linear(512, 512),
nn.ReLU(),
nn.Linear(512, 128),
)
def forward(self, x):
return self.backbone(x)
checkpoint = torch.load(
"simclr_resnet18_baseline_epoch120.pth.tar",
map_location="cpu",
weights_only=True,
)
model = SimCLRResNet18()
model.load_state_dict(checkpoint["state_dict"], strict=True)
model.eval()
```
### Extracting Backbone Features
To use the pretrained ResNet-18 representation without the SimCLR
projection head:
```python
# x should be a preprocessed image tensor with shape [B, 3, 224, 224]
encoder = model.backbone
encoder.fc = nn.Identity()
with torch.no_grad():
features = encoder(x)
print(features.shape)
# torch.Size([1, 512])
```
## Related Models
This checkpoint belongs to the **[Eccentricity-Constrained SimCLR Models (VEDB)](https://hf.co/collections/DM-Diaz/eccentricity-constrained-simclr-models-vedb)** collection.
- [VEDB SimCLR ResNet-18 β Baseline](https://huggingface.co/DM-Diaz/VEDB-SimCLR-ResNet18-Baseline)
- [VEDB SimCLR ResNet-18 β Fovea-Gaze](https://huggingface.co/DM-Diaz/VEDB-SimCLR-ResNet18-Fovea-Gaze)
- [VEDB SimCLR ResNet-18 β Periph](https://huggingface.co/DM-Diaz/VEDB-SimCLR-ResNet18-Periph)
- [VEDB SimCLR ResNet-18 β Periph-NF](https://huggingface.co/DM-Diaz/VEDB-SimCLR-ResNet18-Periph-NF)
## Citation
If you use these model weights in academic work, please cite the associated study:
```bibtex
@inproceedings{diaz2026eccentricity,
author = {Diaz, Dylan M. and Henderson, Margaret M.},
title = {Eccentricity-Constrained CNN Training Reveals Adaptive Information Coding Around the Visual Field},
booktitle = {Proceedings of the 9th Conference on Cognitive Computational Neuroscience},
address = {New York, NY, USA},
year = {2026},
doi = {10.32470/0416gfsq}
}
```
**Proceedings:** Conference on Cognitive Computational Neuroscience 2026
**Preprint:** `arXiv:2607.19316`
### VEDB Citation
Researchers using the underlying Visual Experience Dataset should also cite:
**Greene, M. R., et al. (2024). _The Visual Experience Dataset: Over 200 recorded hours of integrated eye movement, odometry, and egocentric video._ Journal of Vision, 24(11), 6.**
See the [VEDB dataset paper](https://jov.arvojournals.org/article.aspx?articleid=2802101) for the complete author list and citation information.
## License
The model checkpoint in this repository is released under the **Apache License 2.0**.
The VEDB dataset and other third-party resources used in the associated study remain subject to their respective licenses, access requirements, and terms of use. This repository does not redistribute the full VEDB dataset; a small number of example frames are included for illustration of the published visual-field manipulations. |