PoolDINO

Pooling Representation Autoencoders for Efficient Diffusion

Ramón Calvo-González · Youssef Saied · François Fleuret
University of Geneva · Meta

Paper · Code · Project page

PoolDINO compresses pretrained visual representations with a learned spatial pooling operator. It merges each local window of DINOv3 tokens into one token, reducing the sequence length modeled by the diffusion Transformer. The pooling operator is trained jointly with the RGB decoder, preserving the standard two-stage representation-autoencoder training procedure.

This repository contains JAX/Orbax checkpoints for ImageNet-256 RGB reconstruction and class-conditional image generation. These are not Diffusers or Transformers pipeline checkpoints; use the PoolDINO implementation.

Selected generated dog portrait Selected generated motorcycle and rider Selected generated shark underwater Selected generated leather holster

Selected examples. Learned 2 × 2 pooling, 180 training epochs, 100 Euler steps, internal guidance (IG) scale 2.00, no classifier-free guidance (CFG).

How it works

  1. Train the image decoder. A frozen DINOv3-L/16 encoder produces a 16 × 16 grid. A shared affine map pools each non-overlapping window into one 1,024-dimensional token. Tokens are repeated over their original windows before the ViT-XL RGB decoder reconstructs the image. The pooling operator and decoder are trained together.
  2. Train the generator. Freeze the encoder, pooling operator, and RGB decoder. Train a flow-matching Transformer directly on the normalized, compressed latent grid. Internal guidance uses an intermediate generator prediction to guide sampling.

Available checkpoints

All experiment directory names below end in -dinol-vitxl-raev2official-tfds. The historical identifier repeatconv means learned spatial pooling; pool1x1 is the unpooled reference.

Experiment prefix Pooling Unique tokens Token compression Generator epochs Generator step
pool1x1 Unpooled 256 1× 80 100080
repeatconv2x2 Learned 2 × 2 64 4× 80 100080
repeatconv2x2 Learned 2 × 2 64 4× 180 225180
repeatconv2x4 Learned 2 × 4 32 8× 80 100080
repeatconv4x2 Learned 4 × 2 32 8× 80 100080
repeatconv4x4 Learned 4 × 4 16 16× 80 100080

Generator checkpoints are under pooled-generator/<experiment>/<step>/. Their matching RGB decoder checkpoints are under pooled-decoder/<experiment>/40032/, with pooled_latent_stats.npz in the decoder experiment directory.

An additional learned 1 × 1 decoder (repeatconv1x1) is included. It is not the unpooled pool1x1 reference and has no corresponding generator in this release.

The 300-epoch 4 × 4 generator discussed in the paper is not included in the current upload. Neither are the average-pooling generators or the later 222-/390-epoch runs.

Generation results

Reported ImageNet-256 results with 100 Euler steps, 50,000 samples, EMA weights, and IG alone. The scales below reproduce the settings reported in the main results tables, rather than asserting the optimum over every subsequent fine-scale sweep.

Model Epochs IG scale FID ↓ Inception Score ↑
Unpooled reference 80 1.75 1.08 262.00
Learned 2 × 2 80 1.75 1.09 250.36
Learned 2 × 4 80 2.00 1.19 249.58
Learned 4 × 2 80 2.00 1.21 249.96
Learned 4 × 4 80 2.75 1.44 247.07
Learned 2 × 2 180 2.00 1.05 272.00

At 100 sampling steps, 4× and 16× token compression increase latent-sampling throughput by approximately 3.7× and 9.0× relative to the unpooled reference. These measurements use an NVIDIA H100 NVL at batch size 128 and exclude RGB decoding. The 180-epoch run is an extended-training experiment, not a FLOP-matched run.

Download

Install the Hugging Face CLI and authenticate if the repository requires access:

pip install -U huggingface_hub
hf auth login

Download one matching generator, decoder, and latent-statistics set. For example, the 80-epoch 2 × 2 model:

hf download noctrog/pooldino \
  --include 'pooled-decoder/repeatconv2x2-dinol-vitxl-raev2official-tfds/*' \
  --include 'pooled-generator/repeatconv2x2-dinol-vitxl-raev2official-tfds/100080/*' \
  --local-dir checkpoints/pooldino

For the 180-epoch generator, replace 100080 with 225180; it uses the same decoder. Keep the complete checkpoint subtree, including metadata and sharded array files. Downloading a single manifest is not sufficient.

Sampling

The implementation is maintained in the code repository. The intended GPU environment is Linux with CUDA 12 and Python 3.13.

From the code repository, install the locked environment:

uv sync --locked --extra cuda --extra metrics

The following example assumes the download above is in checkpoints/pooldino relative to the code repository. Supply the ImageNet validation-label file described in the code README:

uv run python -m pooldino.eval.gfid_pooled_decoder_adm \
  --generator-path checkpoints/pooldino/pooled-generator/repeatconv2x2-dinol-vitxl-raev2official-tfds \
  --generator-step 100080 \
  --pooled-decoder-path checkpoints/pooldino/pooled-decoder/repeatconv2x2-dinol-vitxl-raev2official-tfds \
  --pooled-decoder-step 40032 \
  --condition-labels-path /path/to/imagenet2012-validation-labels-tfds.npz \
  --protocol raev2_ig --ig-scale 1.75 --generate-only
  • Point --generator-path and --pooled-decoder-path to the experiment directories, not the numeric step directories. Explicitly select the step if multiple checkpoints are downloaded.
  • Keep each decoder's matching pooled_latent_stats.npz. If restoration reports a source-path mismatch after relocation, follow the code README's POOLDINO_ARTIFACT_PATH_MAP instructions; do not disable artifact-identity checks.
  • raev2_ig selects the 100-step evaluation protocol, seed 42, 50,000 images, and BF16 inference. --generate-only skips metric computation, not generation of the full evaluation sample set.
  • IG is active on [0, 0.9] in the implementation's noise-to-data time convention. IG and CFG are neutral at scale 1. The encoder-reconstruction guidance used in appendix ablations has a different scale convention.
  • Preserve the checkpoint's time-shift settings. Check the saved generation_config.txt for the effective sampling protocol.
  • Frozen encoder weights are resolved separately by the implementation. Obtain pretrained weights and datasets under their respective access conditions and terms.

The instructions above are based on the implementation and uploaded checkpoint layout; they are not a claim of a fresh end-to-end GPU reproduction of this Hub download.

Intended use and limitations

These checkpoints support research on representation autoencoders, spatial compression, and class-conditional ImageNet image generation at 256 × 256. They are not text-to-image models.

Compression trades spatial detail and transfer performance for faster generation. Comparable guided generation quality at 4× compression does not imply that all semantic information is preserved: learned pooling gives lower classification-probe accuracy than simpler baselines and does not consistently outperform averaging on segmentation and depth estimation. Stronger compression can produce visible fine-detail artifacts.

Generated images can reflect biases and limitations of ImageNet and the pretrained encoder. The models have not been validated for safety-critical or unrestricted deployment. The selected examples above should not substitute for distribution-level evaluation; the paper includes random sample grids without quality filtering.

Terms and attribution

This model card does not assign a new license to the checkpoints. Upstream components and datasets have their own terms; repository access should not be interpreted as unrestricted commercial-use permission.

PoolDINO builds on the representation-autoencoder framework and RAEv2. The paper and code provide the full methodology, experimental protocols, and upstream citations.

Citation

@misc{calvogonzález2026poolingrepresentationautoencodersefficient,
      title={Pooling Representation Autoencoders for Efficient Diffusion},
      author={Ramón Calvo-González and Youssef Saied and François Fleuret},
      year={2026},
      eprint={2610.09242},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2610.09242},
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for noctrog/pooldino