Spaces:
Paused
Paused
| title: LingBot-Map Streaming 3D Reconstruction | |
| emoji: ๐บ๏ธ | |
| colorFrom: indigo | |
| colorTo: green | |
| sdk: gradio | |
| sdk_version: 5.50.0 | |
| app_file: app.py | |
| pinned: false | |
| license: apache-2.0 | |
| models: | |
| - robbyant/lingbot-map | |
| tags: | |
| - 3d-reconstruction | |
| - point-cloud | |
| - computer-vision | |
| - arxiv:2604.14141 | |
| short_description: Community demo of LingBot-Map streaming 3D reconstruction | |
| # ๐บ๏ธ LingBot-Map โ Streaming 3D Reconstruction | |
| Upload a handful of images or a **short video clip** and get back a navigable 3D point | |
| cloud with the camera trajectory, exported as a GLB you can orbit right in the browser. | |
| ## What this is | |
| A **community** Gradio demo wrapping [LingBot-Map](https://github.com/Robbyant/lingbot-map) | |
| (the [`robbyant/lingbot-map`](https://huggingface.co/robbyant/lingbot-map) checkpoint) โ | |
| **not an official release**. LingBot-Map is a feed-forward 3D foundation model for streaming | |
| reconstruction: it ingests frames one at a time with a paged-KV-cache transformer and | |
| predicts per-frame 3D points, depth, confidence, and camera poses. | |
| ## What this build does carefully | |
| - Uses the model's **own GLB export** (`predictions_to_glb` โ `gr.Model3D`) for a clean, | |
| static point cloud that renders natively in the browser โ no external 3D server to babysit. | |
| - Runs on **ZeroGPU**: FlashInfer is swapped for the built-in **SDPA** attention fallback, | |
| so there's no per-hardware source build to fight. | |
| - Feeds the exporter **world-to-camera** extrinsics (the convention it actually expects), so | |
| the camera frustums and scene orientation come out right. | |
| - Caps the frame count *before* the GPU call and offloads outputs to CPU, so a job always | |
| fits the GPU time budget. | |
| ## Tips for good results | |
| - **Keep it short** โ a few seconds of video, or up to ~48 frames. Longer inputs are capped. | |
| - Frames should **overlap and move smoothly** (a slow pan or orbit). Disjoint, non-overlapping | |
| shots won't reconstruct well. | |
| - **Confidence filter** drops the lowest-confidence *percentage* of points: higher = cleaner | |
| but sparser, lower = denser but noisier. You can re-tune it (and the camera toggle) **without | |
| re-running the model**. | |
| - For ordered image uploads, name frames `000.jpg, 001.jpg, โฆ` โ or just upload a video. | |
| ## Credits & attribution | |
| Built on **LingBot-Map** by the **Robbyant Team** (Apache-2.0). The upstream model builds on | |
| **VGGT** and **DINOv2**; **FlashInfer** was replaced with the SDPA fallback for ZeroGPU | |
| compatibility (an honest, useful note for anyone reproducing this). | |
| - Model: <https://huggingface.co/robbyant/lingbot-map> | |
| - Code: <https://github.com/Robbyant/lingbot-map> | |
| - Paper: <https://arxiv.org/abs/2604.14141> | |
| ```bibtex | |
| @article{chen2026geometric, | |
| title={Geometric Context Transformer for Streaming 3D Reconstruction}, | |
| author={Chen, Lin-Zhuo and Gao, Jian and Chen, Yihang and Cheng, Ka Leong and Sun, Yipengjing and Hu, Liangxiao and Xue, Nan and Zhu, Xing and Shen, Yujun and Yao, Yao and Xu, Yinghao}, | |
| journal={arXiv preprint arXiv:2604.14141}, | |
| year={2026} | |
| } | |
| ``` | |
| ## License | |
| Apache-2.0 โ see `LICENSE.txt`. Upstream copyright headers, including the Meta Platforms | |
| headers inherited from VGGT, are preserved. | |