File size: 4,290 Bytes
13a2b21 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 | ---
license: cc-by-nc-sa-3.0
pipeline_tag: unconditional-image-generation
tags:
- weight-space-learning
- neural-fields
- diffusion
- lora
- shapenet
- ffhq
library_name: pytorch
---
# Weight Space Representation Learning via Neural Field Adaptation — model zoo
Pretrained checkpoints for the CVPR 2026 paper
**Weight Space Representation Learning via Neural Field Adaptation**
(Zhuoqian Yang, Mathieu Salzmann, Sabine Süsstrunk — EPFL).
[Paper](https://arxiv.org/abs/2512.01759) | [Code](https://github.com/IVRL/wsr.pytorch) |
[Project page](https://inrainbws.github.io/wsr/) |
[ShapeNetSDF dataset](https://huggingface.co/datasets/EPFL-IVRL/ShapeNetSDF)
The folder layout mirrors the code repo, so downloading into the repo root restores every
path the configs and `REPRODUCE.md` expect:
```bash
hf download EPFL-IVRL/wsr --local-dir .
```
The pipeline has three stages (Algorithm 1 in the paper). This repo holds each stage's
outputs for **mLoRA-Asym**, the paper's main representation, so any stage can be skipped.
## Stage 1 — base neural fields (3 checkpoints, ~200 MB)
Modulated base fields trained as variational autodecoders. Needed only for the
LoRA / mLoRA representations. `torch.load(...)['g_ema']` is the EMA model used everywhere.
| Path | Dataset |
|---|---|
| `neural_field/train_outputs_2d/nf-res-2d-o32-ffhq-128-5k/checkpoints/checkpoint_500000.pt` | FFHQ 128 (5,000 images) |
| `neural_field/train_outputs_3d/nf-res-3d-airplane/checkpoints/checkpoint_350000.pt` | ShapeNet airplane |
| `neural_field/train_outputs_3d/nf-res-3d-all5k10c-o1/checkpoints/checkpoint_350000.pt` | ShapeNet multi-category (5k10c) |
## Stage 2 — fitted-weight datasets (3 folders, ~3 GB)
One folder per dataset (mLoRA-Asym representation); each holds one PyTorch `state_dict`
per instance (`weights/image_000000.pt` ... in 2D, `weights/shape_000000.pt` ... in 3D).
These are the training data for Stage 3 and the inputs for the reconstruction /
discriminative evaluations (paper Tables 1 and 4).
| Folder | Dataset | Instances |
|---|---|---|
| `neural_field/overfit_outputs_2d/2d_nf-res-2d-o32-ffhq-128-5k_mlora_r4_asym_0.0` | FFHQ 128 | 5,000 |
| `neural_field/overfit_outputs_3d/3d_nf-res-3d-airplane_mlora_r3_asym_0.0` | airplane | 4,045 |
| `neural_field/overfit_outputs_3d/3d_nf-res-3d-all5k10c-o1_mlora_r3_asym_0.0` | multi (10k10c superset) | 10,000 |
Note on the multi-category splits: 5k10c (4,999 shape ids) is a subset of 10k10c
(9,999 ids) but **not** a line-order prefix. The 10,000-instance folder is indexed in
10k10c line order; extracting the 5k10c subset requires the id-to-index mapping from
the split files `ShapeNetSDF/meta/all/{5k10c,10k10c}.txt` in the
[ShapeNetSDF dataset](https://huggingface.co/datasets/EPFL-IVRL/ShapeNetSDF).
## Stage 3 — weight-space diffusion (3 checkpoints, ~16 GB)
One Diffusion Transformer per dataset, trained on the mLoRA-Asym weights (fp32 PyTorch
Lightning checkpoints, weights only, no optimizer state — ~5 GB each). Directory names
equal the Hydra config names in `weight_space_diffusion/configs/diffusion_configs/`;
each holds `last.ckpt`.
| Config (`weight_space_diffusion/checkpoints/<CONFIG>/last.ckpt`) | Dataset | Paper |
|---|---|---|
| `train_ffhq_128_nfres_mlora_r4_asym` | FFHQ 128 | Table 2 |
| `train_airplane_nfres3d_mlora_r3_asym` | ShapeNet airplane | Table 3 |
| `train_all_5k10c_nfres3d_mlora_r3_asym` | ShapeNet multi (5k10c) | Table 3 |
Sampling example (from the code repo, see `REPRODUCE.md` for details):
```bash
cd weight_space_diffusion
python apps/inference.py --config-name=<CONFIG> \
inference.model_path=checkpoints/<CONFIG>/last.ckpt
```
## License
CC BY-NC-SA 3.0. The Stage-3 diffusion code derives from
[HyperDiffusion](https://github.com/Rgtemze/HyperDiffusion) (CC BY-NC-SA 3.0); the
`neural_field` code is MIT with NVIDIA-licensed CUDA ops. See the code repo's `LICENSE`
files and `ACKNOWLEDGEMENT.md`.
## Citation
```bibtex
@inproceedings{yang2026wsr,
title = {Weight Space Representation Learning via Neural Field Adaptation},
author = {Yang, Zhuoqian and Salzmann, Mathieu and S{\"u}sstrunk, Sabine},
booktitle = {Proceedings of the IEEE/CVF Conference on
Computer Vision and Pattern Recognition (CVPR)},
year = {2026}
}
```
|