Instructions to use ruotian/SelectGround-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use ruotian/SelectGround-8B with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3-VL-8B-Instruct") model = PeftModel.from_pretrained(base_model, "ruotian/SelectGround-8B") - Notebooks
- Google Colab
- Kaggle
File size: 5,556 Bytes
cfd843b 7eb63a1 cfd843b 7eb63a1 cfd843b 7eb63a1 cfd843b 7eb63a1 4a027f2 7eb63a1 cfd843b 7eb63a1 cfd843b 7eb63a1 cfd843b 7eb63a1 cfd843b 7eb63a1 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | ---
base_model: Qwen/Qwen3-VL-8B-Instruct
library_name: peft
pipeline_tag: image-text-to-text
license: apache-2.0
datasets:
- ruotian/ContrastGround
tags:
- gui-grounding
- computer-use
- qwen3-vl
- lora
- selectground
---
# SelectGround-8B
SelectGround-8B maps a screenshot and a natural-language instruction to one
click point. This release replaces the earlier ClickContrast-trained checkpoint
with a checkpoint trained from the pinned plain Qwen3-VL-8B-Instruct base on
the `selectground-8b` configuration of
[`ruotian/ContrastGround`](https://huggingface.co/datasets/ruotian/ContrastGround).
It is a single directly trained LoRA checkpoint, not a model soup or weight
aggregate.
## Direct grounding results
| Benchmark | Accuracy | Semantic error |
|---|---:|---:|
| ScreenSpot-Pro | 65.09 | 29.35 |
| UI-Vision | 37.12 | 49.10 |
| OSWorld-G | 69.41 | 20.00 |
UI-Vision is the equal-weight macro over its basic, functional, and spatial
element-grounding subsets. OSWorld-G uses its 510 target-bearing examples;
refusal-only rows are excluded. These public benchmarks were used during model
selection, so results are test-tuned rather than held-out validation estimates.
## Self-Contrastive Grounding
The release also includes Self-Contrastive Grounding, a training-free extension
of the paper's contrast-mining principle. Training mines observed hard
distractors from disagreement between models. At inference, Self-Contrast mines
latent distractors from disagreement between deterministic views of the same
model, then asks every other view to verify each visible coordinate. A proposal
is never scored by the view that generated it, which prevents self-confirmation.
| Inference | ScreenSpot-Pro | UI-Vision | OSWorld-G |
|---|---:|---:|---:|
| Direct | 65.09 | 37.12 | 69.41 |
| Self-Contrast | **71.16** | **44.09** | **72.75** |
The method uses one full-screen view, one 40% incumbent-centered revisit, and
four fixed overlapping 60% views. All crops are enlarged by 2x. Within each
view, coordinate-string mean token log-likelihoods are standardized; evidence
is averaged across non-source views and combined at equal weight with proximity
to the incumbent revisit. This one configuration is shared by all three
benchmarks: there is no benchmark-specific gate, router, prompt, or threshold.
The implementation retains each view's visual prefix after greedy candidate
generation and reuses its KV cache for batched coordinate scoring. It therefore
uses six visual prefills, rather than the twelve prefills of a naive
generate-then-rescore implementation, and requires no weight update.
```bash
python evaluate.py \
--model ruotian/SelectGround-8B \
--benchmark screenspot_pro \
--data data/screenspot-pro \
--self-contrast \
--output outputs/screenspot-pro-self-contrast.jsonl
```
Full 8B ablations, using the same benchmark protocols, are:
| Variant | ScreenSpot-Pro | UI-Vision | OSWorld-G |
|---|---:|---:|---:|
| Full | 71.16 | 44.09 | 72.75 |
| no latent distractors | 71.22 | 43.44 | 69.61 |
| one latent distractor | 70.97 | 43.44 | 70.59 |
| no recurrent anchor | 68.82 | 43.23 | 72.75 |
| no cross-view evidence | 71.16 | 43.46 | 69.41 |
| no anchor proximity | 70.15 | 44.10 | 72.94 |
`self_contrast_manifest.json` records the exact protocol, full counts, split
metrics, artifact checksums, and cache-reuse equivalence test.
## Direct inference
The repository includes the exact loader and evaluator. `visual_merger.pt` must
be loaded in addition to the PEFT adapter; `selectground.py` does this.
```bash
python evaluate.py \
--model ruotian/SelectGround-8B \
--benchmark screenspot_pro \
--data data/screenspot-pro \
--output outputs/screenspot-pro.jsonl
```
Inference uses the full native screenshot, the prompt in `selectground.py`,
Qwen smart resize with `min_pixels=3136` and `max_pixels=8847360`, greedy
decoding for at most 32 tokens, and normalized 0–1000 point coordinates.
## Reproduce training from the plain base
```bash
hf download ruotian/ContrastGround --repo-type dataset \
--local-dir data/ContrastGround
PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True \
accelerate launch --mixed_precision bf16 --num_processes 2 train.py \
--model 8b \
--data data/ContrastGround \
--pairs-file data/ContrastGround/data/selectground-8b/train_pairs.jsonl \
--replay-file data/ContrastGround/data/selectground-8b/train_replays.jsonl \
--output outputs/SelectGround-8B \
--steps 240 --gpus 2 --accumulation 64 \
--learning-rate 3e-5 --selector-learning-rate 1e-4 \
--aux-weight 0.1 --ground-coordinate-weight 1.0 \
--margin 0.3 --pair-weight 0.5 \
--warmup-steps 10 --scheduler-steps 384 \
--holdout-fraction 0.02 --seed 20260819
```
This is SFT coordinate cross-entropy on pair and replay rows plus the paper's
auxiliary competitor-selection loss on pair rows. Pair and replay microbatches
alternate. LoRA uses rank 64, alpha 128, dropout 0.05 on
`q/k/v/o/gate/up/down` projections. The selector reads semantic attention from
layers 18–23. See `training_manifest.json` for the complete recipe and artifact
SHA-256 checksums.
The reference environment used PyTorch 2.11.0+cu128, Transformers 4.57.1,
PEFT 0.19.1, Accelerate 1.13.0, and qwen-vl-utils 0.0.14. CUDA kernels are not
bitwise deterministic; clean runs should be expected to be close rather than
byte-identical.
## License and data
The adapter follows the Apache-2.0 license of the base model. Dataset assets
retain their upstream terms; consult the ContrastGround data card and its
row-level provenance.
|