Spaces:
Running on Zero
Running on Zero
| title: SpatialCLI-8B | |
| emoji: 🧭 | |
| colorFrom: pink | |
| colorTo: blue | |
| sdk: gradio | |
| sdk_version: 6.21.0 | |
| app_file: app.py | |
| short_description: Spatial-reasoning VLM (image+text -> text answer) | |
| python_version: "3.12" | |
| startup_duration_timeout: 1h | |
| # SpatialCLI-8B Demo | |
| A Gradio demo for [SpatialCLI-8B](https://huggingface.co/ZYT-MFM/SpatialCLI-8B), a spatial-reasoning vision-language model fine-tuned from Qwen3-VL-8B-Instruct via the Call-Learn-Internalize framework described in [SpatialCLI: Learning to Reason With Spatial Tools, Then Without Them](https://huggingface.co/papers/2607.27703). | |
| The model internalizes specialist spatial-tool capabilities (localization, segmentation, depth, pose) so it can perform tool-free spatial reasoning directly from an image and a question. | |
| ## Usage | |
| Upload an image and ask a spatial-reasoning question — e.g. "Which object is closer to the camera?", "Describe the depth ordering of the objects", "Where is the red chair located in the image?". | |
| ## Note on weights | |
| `ZYT-MFM/SpatialCLI-8B` weights are being uploaded by the authors. This Space loads them when available and otherwise falls back to the documented base model `Qwen/Qwen3-VL-8B-Instruct`, so the demo continues to function. |