File size: 5,150 Bytes
3b635ee | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 | ---
license: mit
library_name: cod-vae
pipeline_tag: feature-extraction
tags:
- 3d
- shape-reconstruction
- autoencoder
- vae
- occupancy
---
# COD-VAE 16 x 8 (small)
A compact [COD-VAE](https://arxiv.org/abs/2503.08737) that compresses a 3D shape
into **16 latent vectors of 8 dimensions = 128 numbers**, and decodes them back into
an occupancy field. Same latent shape as
[cod-vae-16x8](https://huggingface.co/TimSchneider42/cod-vae-16x8), but a ~4x smaller
network trained for fast decoding: ~35M parameters instead of 188M, with a 14.2M-parameter
decode path instead of 90M.
> **Note:** although the latent shape matches
> [cod-vae-16x8](https://huggingface.co/TimSchneider42/cod-vae-16x8), the two models
> define **different latent spaces** — latents from one cannot be decoded with the
> other.
Trained with [`cod-vae`](https://github.com/TimSchneider42/cod-vae), a PyTorch/JAX
reimplementation of COD-VAE (Cho et al., ICCV 2025). The weights are a self-contained
npz and load with either backend.
## Architecture vs cod-vae-16x8
| | cod-vae-16x8 | this model |
|---|---|---|
| embed dim / heads | 512 / 8 | 256 / 4 |
| encoder | 4 blocks x 3 layers | 3 blocks x 3 layers |
| latent decoder layers | 12 | 6 |
| refinement decoder layers | 12 | 8 |
| total parameters | 188M | ~35M |
| decode-path parameters | 90M | 14.2M |
## Decode speed (H100, float32 + TF32)
| | cod-vae-16x8 | this model |
|---|---|---|
| batch-1 latency (latents -> triplanes) | 4.11 ms | 2.69 ms |
| batch-32 throughput | 2,720 shapes/s | 5,358 shapes/s |
| forward+backward, batch 32 x 2048 queries | 53.0 ms | 25.4 ms |
The dense 128^3 query pass (~3.7 ms) is unchanged; it depends only on the triplane
query head, not on model width or depth.
## Usage
```python
import trimesh
from cod_vae import CODVAE
vae = CODVAE.from_pretrained("TimSchneider42/cod-vae-16x8-small")
mesh = trimesh.load("bunny.obj", force="mesh")
latent, transform = vae.encode_mesh(mesh, return_transform=True) # (16, 8)
reconstruction = vae.decode_mesh(latent, transform=transform) # trimesh.Trimesh
```
Latents can also be computed from raw surface point clouds and decoded at arbitrary
query points:
```python
latents = vae.encode(points) # (N, 3) in [-1, 1]^3
logits = vae.decode(latents, queries) # occupancy logits, positive inside
volume = vae.decode_volume(latents, resolution=128) # dense logit grid
```
Install with `pip install cod-vae[torch,hub]` (or `cod-vae[jax,hub]`).
## Training data
A merged dataset of 110,077 shapes, built with the `cod-vae-dataset` tool:
```bash
cod-vae-dataset data/merged --vecset path/to/shapenet_vecset_root
cod-vae-dataset data/merged \
--hf abc=TimSchneider42/tactile-mnist-abc-dataset-small:0.24435897 --hf-split train \
--num-vol 500000 --num-surface 250000
cod-vae-dataset data/merged \
--hf mnist3d=TimSchneider42/tactile-mnist-mnist3d --hf-split train \
--num-vol 50000 --num-surface 25000
```
| source | shapes | query pools per shape |
|---|---|---|
| ShapeNet (3DShape2VecSet, 55 synsets) | 48,597 | 500k volume + 500k near-surface |
| [tactile-mnist-abc-dataset-small](https://huggingface.co/datasets/TimSchneider42/tactile-mnist-abc-dataset-small) | 50,000 | 500k + 500k |
| [tactile-mnist-mnist3d](https://huggingface.co/datasets/TimSchneider42/tactile-mnist-mnist3d) | 11,480 | 50k + 50k |
Only the training splits are used; the ABC and MNIST3D pool sizes are scaled to the
geometric complexity of each source. Meshes are preprocessed with the original authors'
[sdf_gen](https://github.com/1zb/sdf_gen) recipe.
## Training recipe
Both stages follow the reference schedule (100 + 100 epochs); only the batch layout
differs from cod-vae-16x8 because the smaller model needs fewer GPUs:
| | stage 1 (autoencoder) | stage 2 (latent VAE) |
|---|---|---|
| epochs | 100 | 100 |
| batch | 64 per GPU x 4 GPUs = 256 | 256 per GPU x 2 GPUs = 512 |
| learning rate | 1e-4, scaled by effective batch / 256 | same, halved at epochs 60/70/80/90 |
| dataset repeat | 8 per epoch | 8 per epoch |
| precision | float32 with TF32 matmuls | same |
## Held-out reconstruction quality
| source | held-out shapes | volume IoU | near-surface accuracy |
|---|---|---|---|
| ABC (CAD parts) | 128 | 0.8489 | 0.8085 |
| MNIST3D (embossed digits) | 128 | 0.9105 | 0.8698 |
For reference, the full-size cod-vae-16x8 reaches 0.8733 / 0.8347 on ABC and
0.9231 / 0.8829 on MNIST3D — the size and speed here cost about 0.01–0.03 IoU.
Measured on the test splits of ABC and MNIST3D, which are disjoint from training.
Volume IoU compares `decode(latents, queries) > 0` against ground-truth occupancy on
uniformly sampled query points; near-surface accuracy uses points sampled around the
surface.
## Citation
The model architecture and training recipe are from:
```bibtex
@inproceedings{cho2025cod,
author={Cho, In and Yoo, Youngbeom and Jeon, Subin and Kim, Seon Joo},
title={Representing 3D Shapes with 64 Latent Vectors for 3D Diffusion Models},
booktitle={Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV)},
year={2025}
}
```
|