File size: 3,046 Bytes
6a8e7ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
b447123
6a8e7ea
 
 
 
 
 
 
 
 
 
 
 
 
 
2e31388
 
 
 
 
 
 
 
 
6a8e7ea
 
 
 
 
 
 
 
 
 
 
 
 
 
 
be8a854
6a8e7ea
 
be8a854
6a8e7ea
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
---
license: gpl-3.0
pipeline_tag: image-to-3d
tags:
- novel-view-synthesis
- neural-rendering
- nerf
- 3d
datasets:
- RealEstate10K
- DL3DV-10K
---

# CLiFT: Compressive Light-Field Tokens for Compute-Efficient and Adaptive Neural Rendering

Pretrained checkpoints for **CLiFT** (NeurIPS 2025 spotlight; [arXiv](https://arxiv.org/abs/2507.08776), [project page](https://clift-nvs.github.io/), [code](https://github.com/eric-zqwang/CLiFT)).

CLiFT represents a scene as compressed light-field tokens (CLiFTs): an encoder tokenizes the input views, latent K-means selects centroid tokens, and a neural condenser aggregates each cluster into its centroid. A single trained model supports compute-adaptive rendering — fewer tokens for lower storage/compute, more tokens for higher quality.

## Checkpoints

| File | Description | Place at (in the [code repo](https://github.com/eric-zqwang/CLiFT)) |
|---|---|---|
| `re10k/first_stage.ckpt` | RealEstate10K encoder-decoder (LVSM-style, random token-drop) | `output/re10k_first_stage/training/last.ckpt` |
| `re10k/second_stage.ckpt` | RealEstate10K condenser (full CLiFT model) | `output/re10k_second_stage/training/last.ckpt` |
| `dl3dv/first_stage.ckpt` | DL3DV encoder-decoder, fine-tuned from the RE10K model | `output/dl3dv_first_stage/training/last.ckpt` |
| `dl3dv/second_stage.ckpt` | DL3DV condenser (full CLiFT model) | `output/dl3dv_second_stage/training/last.ckpt` |

The second-stage checkpoints are the full CLiFT models used for the paper's main results. The first-stage checkpoints are used by the random / K-means selection baselines, for K-means annotation, and as the initialization for condenser training.

## Precomputed K-means assignments (for condenser training)

| File | Description |
|---|---|
| `re10k/kmeans_faiss_no_features_merged.zip` | Per-scene K-means assignments for RealEstate10K second-stage training (extract to `re10k_data/kmeans_faiss_no_features_merged/`) |
| `dl3dv/dl3dv_kmeans_faiss_merged.tar.zst` | Per-scene K-means assignments for DL3DV second-stage training (`tar --zstd -xf ... -C Dataset/`) |

Only needed for training the condenser; evaluation just needs the checkpoints.

## Usage

```bash
git clone https://github.com/eric-zqwang/CLiFT.git
cd CLiFT
# download the checkpoints to the paths above, prepare data (see docs/), then e.g.
bash script/eval/eval_clift.sh       # RE10K
bash script/eval/eval_dl3dv.sh 6     # DL3DV, 6 context views
```

See the [test guide](https://github.com/eric-zqwang/CLiFT/blob/main/docs/test.md) for evaluation and the [training guide](https://github.com/eric-zqwang/CLiFT/blob/main/docs/training.md) for the two-stage training pipeline.

## Citation

```bibtex
@inproceedings{Wang2025CLiFT,
  author    = {Wang, Zhengqing and Wu, Yuefan and Chen, Jiacheng and Zhang, Fuyang and Furukawa, Yasutaka},
  title     = {CLiFT: Compressive Light-Field Tokens for Compute Efficient and Adaptive Neural Rendering},
  booktitle = {Advances in Neural Information Processing Systems (NeurIPS)},
  year      = {2025},
}
```