Instructions to use jdzhang0929/Imaginator with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use jdzhang0929/Imaginator with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="jdzhang0929/Imaginator")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("jdzhang0929/Imaginator", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use jdzhang0929/Imaginator with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "jdzhang0929/Imaginator" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jdzhang0929/Imaginator", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/jdzhang0929/Imaginator
- SGLang
How to use jdzhang0929/Imaginator with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "jdzhang0929/Imaginator" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jdzhang0929/Imaginator", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "jdzhang0929/Imaginator" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "jdzhang0929/Imaginator", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use jdzhang0929/Imaginator with Docker Model Runner:
docker model run hf.co/jdzhang0929/Imaginator
| license: apache-2.0 | |
| library_name: transformers | |
| pipeline_tag: image-text-to-text | |
| base_model: | |
| - Qwen/Qwen3-VL-8B-Instruct | |
| tags: | |
| - visual-search | |
| - panoramic | |
| - embodied-ai | |
| - spatial-reasoning | |
| # Imaginator-8B | |
| The Imaginator from *Beyond Thinking: Imagining in 360° for Humanoid Visual | |
| Search*. | |
| An agent searching a 360° scene sees only a narrow field of view at a time. | |
| The Imaginator looks at the views seen so far and says where in the full | |
| panorama the target probably is — including in the parts nobody has looked at | |
| yet. It does not act. A separate, **frozen** search policy decides what to do | |
| with the guess, so the same Imaginator plugs into any policy without retraining | |
| it. | |
| Trained from Qwen3-VL-8B-Instruct. | |
| | Path in this repo | What it is | | |
| |---|---| | |
| | `Imagine-8B/` | this model | | |
| | `HVS-3B/` | the frozen search policy used for the numbers below, from Yu et al. | | |
| ## Input and output | |
| **Input** — every narrow view seen so far, each labelled with the camera | |
| direction it was taken at, plus the instruction: | |
| ``` | |
| [View 1] viewing direction: (0,0) | |
| <image> | |
| [View 2] viewing direction: (0,-30) | |
| <image> | |
| Human Instruction: look for the black and white striped blanket | |
| Decide your next action. | |
| ``` | |
| `viewing direction` is `(yaw, pitch)` in degrees: yaw in `[0,360)` measured | |
| clockwise, pitch in `[-90,90]` with positive up. Views accumulate across the | |
| episode; there is one `<image>` per view, in the order listed. | |
| **Output** — a reasoning block that separates what is visible from what is | |
| inferred, then a single predicted target location: | |
| ``` | |
| <think>[Observed] | |
| - rolled rugs: (342,-4) | |
| - ceiling light: (7,34) | |
| - shopping cart handle: (40,-32) | |
| [Imagined] | |
| - rug sample shelves: (126,-9) | |
| - bedding section: (200,-5) | |
| </think><answer>suggest check(126,-9)</answer> | |
| ``` | |
| `[Observed]` are landmarks the model can see in the given views, with their | |
| absolute panorama coordinates. `[Imagined]` are landmarks it infers lie outside | |
| them — this is the part that carries the spatial prior. The `<answer>` is one | |
| absolute `(yaw, pitch)` guess at where the target is. | |
| The coordinate is a **proposal, not a detection**. On HOS its top-1 hit rate | |
| under the benchmark tolerance is 39.04%, and the system still reaches 62.75, | |
| because the policy is free to reject a bad guess and keep searching. | |
| Downstream, the harness converts `check(yaw,pitch)` into the relative | |
| `rotate(dyaw,dpitch)` or `submit(yaw,pitch)` form the policy was trained on, | |
| depending on whether the target is already within tolerance of the current | |
| view, and appends it to the policy's turn. | |
| ## Results | |
| H\*Bench success rate, from the paper. The policy is identical in both rows and | |
| frozen; the only difference is whether it receives the Imaginator's guess. | |
| | | HOS | HPS | | |
| |---|---|---| | |
| | HVS-3B | 48.04 | 24.12 | | |
| | **HVS-3B + Imaginator-8B** | **62.75** | **39.38** | | |
| ## Usage | |
| vLLM cannot serve a model from a subdirectory of a repo, so fetch it first: | |
| ```bash | |
| hf download jdzhang0929/Imaginator --include "Imagine-8B/*" --local-dir ./checkpoints | |
| vllm serve ./checkpoints/Imagine-8B --port 8001 --served-model-name imaginator | |
| ``` | |
| With transformers the subfolder is addressable directly: | |
| ```python | |
| from transformers import AutoModelForImageTextToText, AutoProcessor | |
| model = AutoModelForImageTextToText.from_pretrained( | |
| "jdzhang0929/Imaginator", subfolder="Imagine-8B", | |
| dtype="auto", device_map="auto") | |
| processor = AutoProcessor.from_pretrained( | |
| "jdzhang0929/Imaginator", subfolder="Imagine-8B") | |
| ``` | |
| Greedy decoding, `max_tokens=4096`. The full two-model loop — view accumulation, | |
| coordinate conversion, and the multi-hypothesis injection used in the paper — | |
| is in the code repo. | |
| ## Training | |
| Two stages. Stage 1 pretrains on 1.92M pseudo-labelled panorama samples | |
| (10,000 steps at batch 192) drawn from Sun360, Matterport3D and DiT360 renders. | |
| Stage 2 is a clean SFT on 4,223 H\*Bench trajectories, 2 epochs at batch 64, | |
| lr 1e-5 cosine. | |
| Stage 2 is reproducible from public data: the trajectories are at | |
| [jdzhang0929/Imagine-in-360-Dataset](https://huggingface.co/datasets/jdzhang0929/Imagine-in-360-Dataset) | |
| and the images they reference ship with H\*Bench. | |
| ## Data separation | |
| Evaluation panoramas were compared against the stage-2 SFT panoramas | |
| exhaustively at the pixel level — 857 × 382 pairs, 36 yaw rotations each, under | |
| two criteria (MAE < 2.0 grey levels for "same photo", Pearson r ≥ 0.90 for | |
| "same viewpoint"). Two overlaps surfaced, both inherited from the original | |
| H\*Bench split, and both are removed from the released SFT trajectories. | |
| Against the 40,453 stage-1 pseudo-label panoramas the same sweep found nothing | |
| at either threshold; the global maximum correlation was 0.8965. | |
| ## Limitations | |
| - English instructions only. | |
| - Top-1 coordinate accuracy is 39.04% (HOS) / 26.12% (HPS). Treat the output as | |
| a prior over where to look, not as a localization result. | |
| - Rotation-only search from a fixed viewpoint; no translation. | |
| - Coordinates assume the equirectangular convention above. A different yaw | |
| origin or pitch sign will silently produce plausible but wrong guesses. | |
| ## Citation | |
| ```bibtex | |
| @article{imagining360, | |
| title = {Beyond Thinking: Imagining in 360{\deg} for Humanoid Visual Search}, | |
| author = {Zhang, Jingdong and others}, | |
| year = {2026} | |
| } | |
| ``` | |
| Built on [H\*Bench](https://huggingface.co/datasets/humanoid-vstar/hstar_bench) | |
| and the HVS models from Yu et al. | |