Update main demo video and prepared cover

#4
by Skywalker0410 - opened
Files changed (4) hide show
  1. README.md +11 -15
  2. assets/demo-poster.jpg +2 -2
  3. assets/demo.mp4 +2 -2
  4. checksums.sha256 +3 -1
README.md CHANGED
@@ -18,20 +18,20 @@ tags:
18
  inference: false
19
  ---
20
 
21
- # <img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/c367acf62cc8255782e3d19bd8c312f9b7ab35db/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything-VLM: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
22
 
23
  **Model family:** [GroundAnything β€” DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) Β· [GroundAnything-VLM β€” autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
24
 
25
  **This repository contains the autoregressive GroundAnything-VLM checkpoint.** The family overview and figures below are shared with the DLM page; use the VLM serving recipe for these weights.
26
 
27
- <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/c367acf62cc8255782e3d19bd8c312f9b7ab35db/assets/fig1-teaser.png" width="100%" alt="GroundAnything: broad visual grounding and parallel visual evidence extraction" /></p>
28
 
29
  ## πŸ”— Quick Links
30
 
31
- - πŸš€ **Online Demo:** Coming soon β€” XXX.
 
32
  - πŸ’» **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything).
33
  - πŸ“„ **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
34
- - πŸ§ͺ **Evaluation data:** Coming soon β€” XXX.
35
 
36
  # Model Overview
37
 
@@ -45,11 +45,11 @@ An optional **self-speculative mode** achieves a **4.51Γ— speedup** over the AR
45
 
46
  ### Demo Videos
47
 
48
- <video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/c367acf62cc8255782e3d19bd8c312f9b7ab35db/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/c367acf62cc8255782e3d19bd8c312f9b7ab35db/assets/demo.mp4"></video>
49
 
50
  **Parallel Decoding**
51
 
52
- <video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/c367acf62cc8255782e3d19bd8c312f9b7ab35db/assets/decoding.mp4"></video>
53
 
54
  ### License/Terms of Use:
55
 
@@ -71,17 +71,13 @@ Global.
71
 
72
  - **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
73
 
74
- ## References(s):
75
-
76
- - [GroundAnything paper and supplementary material](https://arxiv.org/abs/2609.39600).
77
- - [GroundingPI: grounding with visual primitives](https://arxiv.org/abs/2609.39601).
78
 
79
  <details>
80
  <summary>Citation</summary>
81
 
82
  ```bibtex
83
- @misc{yu2026groundanything,
84
- title = {GroundAnything: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
85
  author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
86
  year = {2026},
87
  eprint = {2609.39600},
@@ -103,7 +99,7 @@ Global.
103
  - **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
104
  - **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
105
 
106
- <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/c367acf62cc8255782e3d19bd8c312f9b7ab35db/assets/fig2-architecture.png" width="100%" alt="GroundAnything: vision-language architecture and autoregressive-to-diffusion conversion" /></p>
107
 
108
  ## Input(s):
109
 
@@ -172,7 +168,7 @@ Evaluation code and instructions: [GitHub](https://github.com/groundingpi/Ground
172
 
173
  ## Quantitative Evaluation Benchmarks
174
 
175
- <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/c367acf62cc8255782e3d19bd8c312f9b7ab35db/assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything: GroundAnything and GroundAnything-VLM benchmark overview" /></p>
176
 
177
  ## Inference:
178
 
@@ -277,7 +273,7 @@ After a block is complete, a causal forward reconstructs its authoritative KV ca
277
 
278
  The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
279
 
280
- <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/c367acf62cc8255782e3d19bd8c312f9b7ab35db/assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything: linear and quadratic self-speculative schedules with shared model weights" /></p>
281
 
282
  The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
283
 
 
18
  inference: false
19
  ---
20
 
21
+ # <img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/logo.png" width="40" style="display: inline-block; vertical-align: middle; margin: 0;" alt="GroundAnything logo" /> GroundAnything-VLM: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed
22
 
23
  **Model family:** [GroundAnything β€” DLM / parallel decoding](https://huggingface.co/GroundingPI/GroundAnything) Β· [GroundAnything-VLM β€” autoregressive](https://huggingface.co/GroundingPI/GroundAnything-VLM).
24
 
25
  **This repository contains the autoregressive GroundAnything-VLM checkpoint.** The family overview and figures below are shared with the DLM page; use the VLM serving recipe for these weights.
26
 
27
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/fig1-teaser.png" width="100%" alt="GroundAnything: broad visual grounding and parallel visual evidence extraction" /></p>
28
 
29
  ## πŸ”— Quick Links
30
 
31
+ - πŸš€ **Online Demo:** [Hugging Face Spaces](https://huggingface.co/spaces/GroundingPI/GroundAnything-VLM).
32
+ - 🌐 **Project Page:** [GroundAnything](https://groundingpi.github.io/groundanything/).
33
  - πŸ’» **GitHub Code:** [groundingpi/GroundAnything](https://github.com/groundingpi/GroundAnything).
34
  - πŸ“„ **Paper:** [arXiv:2609.39600](https://arxiv.org/abs/2609.39600).
 
35
 
36
  # Model Overview
37
 
 
45
 
46
  ### Demo Videos
47
 
48
+ <video controls playsinline preload="none" width="100%" poster="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/demo-poster.jpg" src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/demo.mp4"></video>
49
 
50
  **Parallel Decoding**
51
 
52
+ <video controls playsinline preload="none" width="100%" src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/decoding.mp4"></video>
53
 
54
  ### License/Terms of Use:
55
 
 
71
 
72
  - **Paper [09/30/2026]:** [GroundAnything](https://arxiv.org/abs/2609.39600).
73
 
 
 
 
 
74
 
75
  <details>
76
  <summary>Citation</summary>
77
 
78
  ```bibtex
79
+ @misc{yu2026groundanythingreconcilingparalleldecoding,
80
+ title = {{GroundAnything}: Reconciling Parallel Decoding with Precise Visual Grounding at Flash Speed},
81
  author = {Qize Yu and Lianrui Fan and Bowen Ping and Xini Ding and Zetian Song and Junbo Niu and Kaixuan Wang and Tianxing Chen and Yue Chen and Minghua He and Yuran Wang and Jie Huang and Haojun Zhang and Min Chen and Hao Li and Wenxuan Song and Ruihai Wu and Xianming Liu and Shilong Liu and Shuchang Zhou and Ping Luo and Shiyu Huang},
82
  year = {2026},
83
  eprint = {2609.39600},
 
99
  - **Spatial vocabulary:** 1,000 coordinate tokens shared with semantic labels and protocol markers.
100
  - **DLM conversion:** the shared decoder and vocabulary head support both causal prediction and bidirectional response-block denoising. A mask token is added for diffusion generation.
101
 
102
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/fig2-architecture.png" width="100%" alt="GroundAnything: vision-language architecture and autoregressive-to-diffusion conversion" /></p>
103
 
104
  ## Input(s):
105
 
 
168
 
169
  ## Quantitative Evaluation Benchmarks
170
 
171
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/fig7-grounding-performance.png" width="100%" alt="GroundAnything: GroundAnything and GroundAnything-VLM benchmark overview" /></p>
172
 
173
  ## Inference:
174
 
 
273
 
274
  The model uses its own shared weights to draft tokens with bidirectional attention and verify them with causal attention. Verification accepts the **longest consecutive matching prefix**, stops at the first mismatch, applies the causal correction, and discards the rejected suffix cache states. The shipped speculative route uses greedy verification; it is not a general stochastic speculative sampler.
275
 
276
+ <p align="center"><img src="https://huggingface.co/GroundingPI/GroundAnything-VLM/resolve/61cf85b86dc3926b258ea9a045504b2752b896fd/assets/fig6-self-speculative-decoding.png" width="100%" alt="GroundAnything: linear and quadratic self-speculative schedules with shared model weights" /></p>
277
 
278
  The documented `--decoder speculative` service is the linear shared-weight route. Exact greedy verification is relative to the converted model's causal branch; it does not imply identical outputs to the separately trained GroundAnything-VLM checkpoint.
279
 
assets/demo-poster.jpg CHANGED

Git LFS Details

  • SHA256: 288c6121ec41e3114fe98221cf795381b1aace7c99660da945164be88af331f0
  • Pointer size: 131 Bytes
  • Size of remote file: 217 kB

Git LFS Details

  • SHA256: 3c282238bb1598093c078e57dcc1db4b3ce8fb1f39db6c09ca90e0686f1043a2
  • Pointer size: 131 Bytes
  • Size of remote file: 725 kB
assets/demo.mp4 CHANGED
@@ -1,3 +1,3 @@
1
  version https://git-lfs.github.com/spec/v1
2
- oid sha256:1d987aac49b390c0411faecc33892873fb0b8b811ff1f961d8a87cb39c406285
3
- size 22955151
 
1
  version https://git-lfs.github.com/spec/v1
2
+ oid sha256:cddaec9f528f9ebf5ac627850eefcb3f14ee55e168776a1b650f91353ce8293e
3
+ size 45075106
checksums.sha256 CHANGED
@@ -1,5 +1,5 @@
1
  c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
2
- c81ef575787c65a89274543f261edde048e5512b526a089e2137c0d03db67966 README.md
3
  7316e325dde0dd407bcc80ed5d3d080223e54572a4685b2aa1ee52041e41d466 added_tokens.json
4
  a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
5
  2a23a7cf76128488c8866582bab06ffb1db387f7f7de82e06fafea65514ce10d config.json
@@ -24,3 +24,5 @@ d9089e4621f850e727a86c64110730d10feb459a9532e8bd8b77c8303d417bd3 tokenizer_conf
24
  ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910 vocab.json
25
  cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE-Apache-2.0
26
  20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE-Kimi-K3
 
 
 
1
  c419c191a33774f4e1f834b3fb36033dc84e8ba316a3d1935eed1fa529884a74 LICENSE
2
+ 201de2c582a112ca5daab81aea0803538334ee90d1c9b793d7942d98019bee50 README.md
3
  7316e325dde0dd407bcc80ed5d3d080223e54572a4685b2aa1ee52041e41d466 added_tokens.json
4
  a0bc6f6fc7a29a80017a433e8f03a1cc1236e838a944a2d034295a60c4f2fddb chat_template.jinja
5
  2a23a7cf76128488c8866582bab06ffb1db387f7f7de82e06fafea65514ce10d config.json
 
24
  ca10d7e9fb3ed18575dd1e277a2579c16d108e32f27439684afa0e10b1440910 vocab.json
25
  cfc7749b96f63bd31c3c42b5c471bf756814053e847c10f3eb003417bc523d30 LICENSE-Apache-2.0
26
  20c797ce19af0c17de52c6afb144644768a591c521655f5ebf5712c9850f2887 LICENSE-Kimi-K3
27
+ cddaec9f528f9ebf5ac627850eefcb3f14ee55e168776a1b650f91353ce8293e assets/demo.mp4
28
+ 3c282238bb1598093c078e57dcc1db4b3ce8fb1f39db6c09ca90e0686f1043a2 assets/demo-poster.jpg