GUI-Primitives: Diagnosing Spatial Reasoning Failures in Vision-Language GUI Grounding
Abstract
A benchmark of contrastive GUI instructions reveals that vision-language models mostly fail to localize interface elements rather than misunderstand spatial relations, though marking candidates substantially improves selection.
Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether models bind relational language to the correct element. We introduce GUI-Primitives, a 994-item benchmark of contrastive instruction pairs over seven spatial relations in graphical user interfaces (left/right, above/below, containment, alignment, proximity, list ordinal, occlusion). Each pair holds the screenshot and anchor fixed while changing the relation expression, so the correct target moves between two designated candidates. Five annotators validate a 196-item subset (κ= 0.94 well-formedness; κ= 0.79 target selection). Nineteen vision-language models reach at most 32% strict point-in-box accuracy. Because models emit unconstrained coordinates, we classify each prediction by the candidate region it falls within. Predictions fall outside both candidates on 60-92% of items. Conditional on falling within a candidate region, target selection reaches 0.82-0.90 for horizontal position, vertical position, proximity, and list ordinal, but does not differ significantly from 0.50 for containment and occlusion: most failures reflect candidate localization rather than relation understanding. Across ten models, benchmark accuracy correlates with ScreenSpot-Pro accuracy (Spearman ρ= +0.74), an exploratory association at this sample size. Marking the two designated candidates raises selection accuracy by 35--57 percentage points, an oracle diagnostic that supplies the candidate set rather than a deployable method. We release the benchmark, predictions, and code.
Community
EMNLP 2026 Main (CORE A; 15.4% acceptance). GUI-PRIMITIVES is a 994-item benchmark diagnosing spatial reasoning failures in GUI grounding across 19 vision-language models and seven spatial primitives.
This is an automated message from the Librarian Bot. I found the following papers similar to this paper.
The following papers were recommended by the Semantic Scholar API
- GUI-Lens: Coarse-to-Fine Cropping for GUI Grounding with General-Purpose VLMs (2026)
- MindEdit-Bench: Benchmarking Object-Level Counterfactual Spatial Reasoning in VLMs from In-the-Wild Photos (2026)
- Grounding Isn't Knowing: Do VLMs Need Object Localization for Spatial Reasoning? (2026)
- SEER: A Self-Grounded Evidence Interface for Controlled Spatial Relation Classification (2026)
- ChartProbe: A Diagnostic Study on Visual Reasoning through Perception, Grounding, and Simple Reasoning (2026)
- SkillLens: Visual Skill Cards for Retrieval-Augmented GUI Action Prediction and On-Policy Distillation (2026)
- PointRL: Learning Point-Level Vision-Language Grounding from Verifiable Annotation Evidence (2026)
Please give a thumbs up to this comment if you found it helpful!
If you want recommendations for any Paper on Hugging Face checkout this Space
You can directly ask Librarian Bot for paper recommendations by tagging it in a comment: @librarian-bot recommend
Get this paper in your agent:
hf papers read 2608.21832 Don't have the latest CLI?
curl -LsSf https://hf.co/cli/install.sh | bash Models citing this paper 0
No model linking this paper
Datasets citing this paper 0
No dataset linking this paper
Spaces citing this paper 0
No Space linking this paper
Collections including this paper 0
No Collection including this paper