| --- |
| license: other |
| pipeline_tag: image-to-3d |
| tags: |
| - 3d-scene-reconstruction |
| - rgb-d |
| - single-image |
| - flow-matching |
| - textured-mesh |
| --- |
| |
| # FIRE3D: Feed-forward Interactive 3D Scene Reconstruction Within A Minute |
|
|
| [Paper](https://arxiv.org/pdf/2609.08848) | |
| [Code](https://github.com/xiahongchi/Fire3D) | |
| [Inference data](https://huggingface.co/datasets/hongchi/Fire3D) |
|
|
| Fire3D is a unified feed-forward framework that transforms a single RGB image |
| or casual RGB video into simulation-ready 3D scene assets. It predicts a |
| compositional scene representation with object-level 6-DoF pose, bounding box, |
| mesh geometry, and texture, without test-time optimization. |
|
|
| This release provides the inference code, model checkpoints, processed example |
| inputs, and frozen protocols needed to reproduce Fire3D results on iTHOR, |
| Imaginarium, ScanNet++, and single-image scenes. |
|
|
| The inference code and installation instructions are maintained at |
| [xiahongchi/Fire3D](https://github.com/xiahongchi/Fire3D). Processed release |
| inputs are hosted at |
| [datasets/hongchi/Fire3D](https://huggingface.co/datasets/hongchi/Fire3D). |
|
|
| ## Overview |
|
|
|  |
|
|
| **Teaser.** Fire3D reconstructs a compositional 3D scene from a single RGB |
| image or casual RGB video. The output contains object-level pose, geometry, |
| and material assets that can be rendered, edited, and used in simulation. |
|
|
| [](assets/method_overview.pdf) |
|
|
| **Method overview.** Fire3D lifts RGB-D observations into a shared 3D scene |
| representation, predicts object instances and 6-DoF poses, and reconstructs |
| the sparse structure, shape, and PBR fields with three cascaded flow-matching |
| models. Batched decoders and mesh post-processing produce the final textured |
| scene assets. |
|
|
| [](assets/hcvae.pdf) |
|
|
| **HC-VAE.** The Hierarchical Compression VAE reduces each sparse SC-VAE shape |
| or material field from a `32^3 x 32` representation to an `8^3 x 64` latent. |
| This 32x reduction makes scene-level batched flow sampling practical while |
| retaining the frozen TRELLIS.2 decoding path. |
|
|
| ## Model Bundle |
|
|
| The repository contains the scene perception checkpoint, the three |
| flow-matching checkpoints, sparse VAE encoders and decoders, latent normalization |
| statistics, and the TRELLIS.2 shape/PBR decoders used by the frozen release |
| protocols. `manifest.json` records the byte size and SHA-256 digest of every |
| inference file; `checksums.sha256` provides the same values in standard form. |
|
|
| Fire3D's Hierarchical Compression VAE (HC-VAE) compresses each sparse SC-VAE |
| shape and material field from a `32^3` representation with 32 channels into an |
| `8^3` representation with 64 channels. This 32x reduction enables batched flow |
| sampling and VAE decoding across many scene instances while retaining the |
| frozen TRELLIS.2 SC-VAE decoding path. |
|
|
| | Component | Bundle path | |
| |---|---| |
| | Scene perception | `perception/model.pt` | |
| | Sparse-structure flow | `reconstruction/flows/ss/model.pt` | |
| | Shape flow | `reconstruction/flows/shape/model.pt` | |
| | PBR flow | `reconstruction/flows/pbr/model.pt` | |
| | Sparse-structure VAE | `reconstruction/vae/ss/ckpts/encoder.pt`, `decoder.pt` | |
| | Shape HC-VAE | `reconstruction/vae/shape/ckpts/encoder.pt`, `decoder.pt` | |
| | PBR HC-VAE | `reconstruction/vae/pbr/ckpts/encoder.pt`, `decoder.pt` | |
| | DINOv3 encoder | `external/dinov3_vitl16_pretrain_lvd1689m-8aa4cbdd.pth` | |
| | TRELLIS.2 shape decoder | `external/trellis2/shape_dec_next_dc_f16c32_fp16.safetensors` | |
| | TRELLIS.2 PBR decoder | `external/trellis2/tex_dec_next_dc_f16c32_fp16.safetensors` | |
|
|
| Public filenames intentionally do not encode private training iteration |
| numbers. |
|
|
| The sparse-structure VAE encoder and decoder operate on binary `8^3` object |
| supports through an `8 x 2 x 2 x 2` latent. The normal joint-training recipe |
| uses BCE-with-logits plus KL and is available in the |
| [code release](https://github.com/xiahongchi/Fire3D/blob/main/training/README.md#sparse-structure-vae). |
| When training a new SS flow, regenerate its latent cache and statistics with a |
| matched VAE encoder. The released pretrained flow remains pinned to its own |
| latent artifact contract. |
|
|
| The root `config.json` is the model repository's Hugging Face query file. A |
| normal `snapshot_download` includes it in the model snapshot, allowing the Hub |
| to maintain the standard repository download count without client-side |
| telemetry. See the [Hugging Face download-statistics |
| documentation](https://huggingface.co/docs/hub/models-download-stats). |
|
|
| The default protocols use classifier-free guidance 3, flow and VAE batches of |
| 16, and the batch CuMesh postprocessor. The iTHOR, Imaginarium, and ScanNet++ |
| protocols enable the validated background room-box fit before background |
| reconstruction. |
|
|
| The released HC-VAE recipe uses one process on one GPU; its configured |
| per-GPU batch size is therefore the effective global batch size. |
|
|
| ## Training Data |
|
|
| Fire3D does not redistribute its training corpora. All source datasets are |
| publicly available and retain their original licenses and access terms. For |
| object training data, use |
| [TRELLIS-500K](https://huggingface.co/datasets/JeffreyXiang/TRELLIS-500K), |
| which provides the metadata and preparation tools for the ObjaverseXL, ABO, |
| 3D-FUTURE, and HSSD sources used by the object pipeline. Official scene-data |
| links and Fire3D preprocessing entry points are listed in the |
| [code release](https://github.com/xiahongchi/Fire3D/tree/main/data_processing). |
|
|
| ## Use |
|
|
| ```bash |
| git clone https://github.com/xiahongchi/Fire3D.git |
| cd Fire3D |
| bash scripts/install.sh |
| conda activate fire3d |
| bash scripts/install_blender.sh |
| |
| fire3d download --models |
| fire3d download --data --dataset ithor --scene-id iTHOR_FloorPlan312_physics |
| bash scripts/run_ithor_example.sh |
| ``` |
|
|
| The downloader places this repository under `checkpoints/Fire3D/` and verifies |
| the release manifest before inference. |
|
|
| ## Scope |
|
|
| The released checkpoints are intended for inference through the frozen Fire3D |
| protocols. Performance can degrade for scenes with inaccurate camera poses or |
| depth, severe occlusion, unusual scale, non-room backgrounds, or objects far |
| outside the training distribution. The generated meshes should be reviewed |
| before use in safety-critical or physically deployed applications. |
|
|
| ## Licenses |
|
|
| Original Fire3D code is MIT licensed. This model bundle includes DINOv3 and |
| selected TRELLIS.2 files, which retain their upstream licenses. Review the |
| included `licenses/` directory and the repository's |
| `THIRD_PARTY_NOTICES.md` before use or redistribution. |
|
|
| ## Citation |
|
|
| ```bibtex |
| @article{xia2026fire3d, |
| title={{FIRE3D}: Feed-forward Interactive 3D Scene Reconstruction Within A Minute}, |
| author={Xia, Hongchi and Cheng, Tianhang and Ma, Wei-Chiu and Wang, Shenlong}, |
| journal={arXiv preprint arXiv:2609.08848}, |
| year={2026}, |
| url={https://arxiv.org/pdf/2609.08848} |
| } |
| ``` |
|
|