ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes
Abstract
ZipTok3D organizes 3D geometry into compact global-token prefixes with iterative decoding to achieve high-fidelity reconstruction from extremely short token sequences.
Compact token sequences are essential for efficient 3D generation. However, existing 3D tokenizers typically organize latent representations either over spatial regions or as fixed-size sets of global tokens, both suffering sharp reconstruction degradation when compressed to extremely low token budgets. In this paper, we present ZipTok3D, a 3D tokenizer designed for high-fidelity reconstruction from extremely short token sequences. Its key idea is to organize object geometry into progressively informative global-token prefixes and unfold these compact representations through iterative decoding. Specifically, nested dropout randomly truncates the latent sequence after encoding during training and requires each retained prefix to reconstruct the complete object, thereby prioritizing essential geometric information in the leading tokens. The decoder then repeatedly applies a parameter-shared Transformer block to recover fine-grained geometry from each prefix without a separate generative sampling stage. With the same token dimension, ZipTok3D achieves reconstruction quality comparable to the 32-token COD-VAE baseline using only one token on ShapeNet and four on TRELLIS, yielding 32times and 8times shorter token sequences, respectively.
Community
ZipTok3D: High-Fidelity 3D Tokenization with Compact Token Prefixes
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- Keep-or-Drop? Adaptive Tokenizer for Compact Video Representation (2026)
- GenSplatCodec: Feed-Forward Gaussian Splatting Compression via One-Step Diffusion (2026)
- V-RAE: Rethinking Video Latent Spaces for Generation (2026)
- Beyond Pixels: From Video Priors to 4D Worlds (2026)
- CoANeRV: Coordinate-Aware Token-Space Neural Video Representation (2026)
- Visual Token Codec: Unleashing Spatial Redundancy for ViT Feature Coding (2026)
- CuteTTS: Efficient and High-Quality Speech Synthesis via Autoregressive Modeling of Continuous Latents (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2609.01740 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper