Image-Text-to-Text
Transformers
Safetensors
English
Chinese
groundinganything_vlm
text-generation
visual-grounding
object-detection
referring-expression-comprehension
pointing
ocr
document-layout
custom-code
conversational
custom_code
Instructions to use GroundingPI/GroundAnything-VLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use GroundingPI/GroundAnything-VLM with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="GroundingPI/GroundAnything-VLM", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("GroundingPI/GroundAnything-VLM", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use GroundingPI/GroundAnything-VLM with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "GroundingPI/GroundAnything-VLM" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything-VLM", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/GroundingPI/GroundAnything-VLM
- SGLang
How to use GroundingPI/GroundAnything-VLM with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything-VLM" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything-VLM", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "GroundingPI/GroundAnything-VLM" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "GroundingPI/GroundAnything-VLM", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use GroundingPI/GroundAnything-VLM with Docker Model Runner:
docker model run hf.co/GroundingPI/GroundAnything-VLM
Update main demo video and prepared cover
#4
by Skywalker0410 - opened
- README.md +11 -15
- assets/demo-poster.jpg +2 -2
- assets/demo.mp4 +2 -2
- checksums.sha256 +3 -1
README.md
CHANGED
|
@@ -18,20 +18,20 @@ tags:
|
|
| 18 |
inference: false
|
| 19 |
---
|
| 20 |
|
| 21 |
-
# <img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/
|
| 22 |
|
| 23 |
**Model family:** [GroundAnything β DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) Β· [GroundAnything-VLM β autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
|
| 24 |
|
| 25 |
**This repository contains the autoregressive GroundAnything-VLM checkpoint.** The family overview and figures below are shared with the DLM page; use the VLM serving recipe for these weights.
|
| 26 |
|
| 27 |
-
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/
|
| 28 |
|
| 29 |
## π Quick Links
|
| 30 |
|
| 31 |
-
- π **Online Demo:**
|
|
|
|
| 32 |
- π» **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything).
|
| 33 |
- π **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
|
| 34 |
-
- π§ͺ **Evaluation data:** Coming soon β XXX.
|
| 35 |
|
| 36 |
# Model Overview
|
| 37 |
|
|
@@ -45,11 +45,11 @@ An optional **self-speculative mode** achieves a **4.51Γ speedup** over the AR
|
|
| 45 |
|
| 46 |
### Demo Videos
|
| 47 |
|
| 48 |
-
<video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/
|
| 49 |
|
| 50 |
**Parallel Decoding**
|
| 51 |
|
| 52 |
-
<video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/
|
| 53 |
|
| 54 |
### License/Terms of Use:
|
| 55 |
|
|
@@ -71,17 +71,13 @@ Global.
|
|
| 71 |
|
| 72 |
- **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
|
| 73 |
|
| 74 |
-
## References(s):
|
| 75 |
-
|
| 76 |
-
- [GroundAnything paper and supplementary material](https://arxiv.org/abs/2609.39600).
|
| 77 |
-
- [GroundingPI: grounding with visual primitives](https://arxiv.org/abs/2609.39601).
|
| 78 |
|
| 79 |
<details>
|
| 80 |
<summary>Citation</summary>
|
| 81 |
|
| 82 |
```bibtex
|
| 83 |
-
@misc{
|
| 84 |
-
title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
|
| 85 |
author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
|
| 86 |
year = {2026},
|
| 87 |
eprint = {2609.39600},
|
|
@@ -103,7 +99,7 @@ Global.
|
|
| 103 |
- **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
|
| 104 |
- **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
|
| 105 |
|
| 106 |
-
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/
|
| 107 |
|
| 108 |
## Input(s):
|
| 109 |
|
|
@@ -172,7 +168,7 @@ Evaluation code and instructions: [GitHub](https://github.com/groundingpi/Ground
|
|
| 172 |
|
| 173 |
## Quantitative Evaluation Benchmarks
|
| 174 |
|
| 175 |
-
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/
|
| 176 |
|
| 177 |
## Inference:
|
| 178 |
|
|
@@ -277,7 +273,7 @@ After a block is complete, a causal forward reconstructs its authoritative KV ca
|
|
| 277 |
|
| 278 |
The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
|
| 279 |
|
| 280 |
-
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/
|
| 281 |
|
| 282 |
The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
|
| 283 |
|
|
|
|
| 18 |
inference: false
|
| 19 |
---
|
| 20 |
|
| 21 |
+
# <img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything-VLM: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
|
| 22 |
|
| 23 |
**Model family:** [GroundAnything β DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) Β· [GroundAnything-VLM β autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
|
| 24 |
|
| 25 |
**This repository contains the autoregressive GroundAnything-VLM checkpoint.** The family overview and figures below are shared with the DLM page; use the VLM serving recipe for these weights.
|
| 26 |
|
| 27 |
+
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/fig1-teaser.png" width="100%" alt="GroundAnything: broad visual grounding and parallel visual evidence extraction" /></p>
|
| 28 |
|
| 29 |
## π Quick Links
|
| 30 |
|
| 31 |
+
- π **Online Demo:** [Hugging Face Spaces](https://huggingface.co/spaces/GroundingPI/GroundAnything-VLM).
|
| 32 |
+
- π **Project Page:** [GroundAnything](https://groundingpi.github.io/groundanything/).
|
| 33 |
- π» **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything).
|
| 34 |
- π **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
|
|
|
|
| 35 |
|
| 36 |
# Model Overview
|
| 37 |
|
|
|
|
| 45 |
|
| 46 |
### Demo Videos
|
| 47 |
|
| 48 |
+
<video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/demo.mp4"></video>
|
| 49 |
|
| 50 |
**Parallel Decoding**
|
| 51 |
|
| 52 |
+
<video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/decoding.mp4"></video>
|
| 53 |
|
| 54 |
### License/Terms of Use:
|
| 55 |
|
|
|
|
| 71 |
|
| 72 |
- **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
|
| 73 |
|
|
|
|
|
|
|
|
|
|
|
|
|
| 74 |
|
| 75 |
<details>
|
| 76 |
<summary>Citation</summary>
|
| 77 |
|
| 78 |
```bibtex
|
| 79 |
+
@misc{yu2026groundanythingreconcilingparalleldecoding,
|
| 80 |
+
title = {{GroundAnything}: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
|
| 81 |
author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
|
| 82 |
year = {2026},
|
| 83 |
eprint = {2609.39600},
|
|
|
|
| 99 |
- **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
|
| 100 |
- **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
|
| 101 |
|
| 102 |
+
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/fig2-architecture.png" width="100%" alt="GroundAnything: vision-language architecture and autoregressive-to-diffusion conversion" /></p>
|
| 103 |
|
| 104 |
## Input(s):
|
| 105 |
|
|
|
|
| 168 |
|
| 169 |
## Quantitative Evaluation Benchmarks
|
| 170 |
|
| 171 |
+
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything: GroundAnything and GroundAnything-VLM benchmark overview" /></p>
|
| 172 |
|
| 173 |
## Inference:
|
| 174 |
|
|
|
|
| 273 |
|
| 274 |
The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
|
| 275 |
|
| 276 |
+
<p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything: linear and quadratic self-speculative schedules with shared model weights" /></p>
|
| 277 |
|
| 278 |
The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
|
| 279 |
|
assets/demo-poster.jpg
CHANGED
|
Git LFS Details
|
|
Git LFS Details
|
assets/demo.mp4
CHANGED
|
@@ -1,3 +1,3 @@
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
-
oid sha256:
|
| 3 |
-
size
|
|
|
|
| 1 |
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:cddaec9f528f9ebf5ac627850eefcb3f14ee55e168776a1b650f91353ce8293e
|
| 3 |
+
size 45075106
|
checksums.sha256
CHANGED
|
@@ -1,5 +1,5 @@
|
|
| 1 |
c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
|
| 2 |
-
|
| 3 |
7316e325dde0dd407bcc80ed5d3d080223e54572a4685b2aa1ee52041e41d466 added_tokens.json
|
| 4 |
a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
|
| 5 |
2a23a7cf76128488c8866582bab06ffb1db387f7f7de82e06fafea65514ce10d config.json
|
|
@@ -24,3 +24,5 @@ d9089e4621f850e727a86c64110730d10feb459a9532e8bd8b77c8303d417bd3 tokenizer_conf
|
|
| 24 |
ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910 vocab.json
|
| 25 |
cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE-Apache-2.0
|
| 26 |
20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE-Kimi-K3
|
|
|
|
|
|
|
|
|
| 1 |
c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
|
| 2 |
+
201de2c582a112ca5daab81aea0803538334ee90d1c9b793d7942d98019bee50 README.md
|
| 3 |
7316e325dde0dd407bcc80ed5d3d080223e54572a4685b2aa1ee52041e41d466 added_tokens.json
|
| 4 |
a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
|
| 5 |
2a23a7cf76128488c8866582bab06ffb1db387f7f7de82e06fafea65514ce10d config.json
|
|
|
|
| 24 |
ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910 vocab.json
|
| 25 |
cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE-Apache-2.0
|
| 26 |
20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE-Kimi-K3
|
| 27 |
+
cddaec9f528f9ebf5ac627850eefcb3f14ee55e168776a1b650f91353ce8293e assets/demo.mp4
|
| 28 |
+
3c282238bb1598093c078e57dcc1db4b3ce8fb1f39db6c09ca90e0686f1043a2 assets/demo-poster.jpg
|