File size: 7,986 Bytes
f9a59c1
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
---
pipeline_tag: image-text-to-text
tags:
  - multimodal
  - vision-language
  - gui-agent
  - diffusion-language-model
  - mixture-of-experts
  - block-diffusion
---

<h1 align="center">LLaDA-UI</h1>

<h3 align="center">Bringing Block-wise Diffusion to Vision-Language GUI Agents</h3>

<p align="center">
  <a href="https://arxiv.org/abs/2609.13287"><img src="https://img.shields.io/badge/Paper-arXiv-B31B1B?style=for-the-badge&logo=arxiv&logoColor=white" alt="Paper"></a>
  <a href="https://www.inclusion-ai.org/LLaDA-UI/"><img src="https://img.shields.io/badge/Project-Page-4C8BF5?style=for-the-badge&logo=googlechrome&logoColor=white" alt="Project Page"></a>
  <a href="https://huggingface.co/inclusionAI/LLaDA-UI"><img src="https://img.shields.io/badge/%F0%9F%A4%97-Model-FFD21E?style=for-the-badge" alt="Hugging Face Model"></a>
  <a href="https://github.com/inclusionAI/LLaDA-UI"><img src="https://img.shields.io/badge/GitHub-Code-181717?style=for-the-badge&logo=github" alt="GitHub Code"></a>
</p>

<p align="center">
  <a href="#highlights"><img src="https://img.shields.io/badge/Diffusion-Model-7B61FF" alt="Diffusion Model"></a>
  <a href="#highlights"><img src="https://img.shields.io/badge/Multimodal-LLM-00A67E" alt="Multimodal LLM"></a>
  <a href="#highlights"><img src="https://img.shields.io/badge/GUI-Agent-EA4AAA" alt="GUI Agent"></a>
  <a href="#highlights"><img src="https://img.shields.io/badge/Mixture--of--Experts-MoE-2F80ED" alt="Mixture of Experts"></a>
</p>

LLaDA-UI is an MoE-based, block-wise diffusion vision-language GUI agent. It
understands screenshots at their native aspect ratio and produces grounded
coordinates or structured actions for mobile, desktop, and web interfaces.

This repository contains the model weights. The inference code, example
requests, and SGLang serving instructions are available in the
[LLaDA-UI code repository](https://github.com/inclusionAI/LLaDA-UI).

<p align="center">
  <img src="assets/figure1_overview_animated.gif" width="100%" alt="LLaDA-UI GUI-agent benchmark comparison and animated qualitative diffusion decoding">
</p>

<p align="center"><em>Figure 1. LLaDA-UI GUI-agent performance and qualitative block-wise diffusion decoding. The radar chart and GUI observation remain static while the model output is progressively denoised.</em></p>

## Highlights

- **Diffusion-native GUI agent:** reasons about screenshots and generates GUI
  actions with block-wise diffusion decoding rather than autoregressive
  token-by-token decoding.
- **Mixture-of-Experts architecture:** approximately 16.7B total parameters.
- **Native-resolution vision:** preserves the aspect ratio and visual details
  of mobile, desktop, and web screenshots.
- **Cross-platform interaction:** supports GUI grounding and agent tasks on
  mobile, desktop, and web interfaces.
- **Structured output:** emits normalized coordinates or tagged reasoning and
  executable actions.

## Model Details

| Item | Description |
|---|---|
| Model type | MoE block-wise diffusion vision-language GUI agent |
| Total parameters | Approximately 16.7B |
| Language backbone | LLaDA2.0-mini-base |
| Vision encoder | Native-resolution ViT initialized from SigLIP, with 2D RoPE |
| Output | Grounded coordinates, text reasoning, and structured GUI actions |
| Coordinate convention | Integer coordinates normalized to `[0, 999]` |
| Supported domains | GUI grounding, mobile, desktop, and web |
| Weight format | BF16 Safetensors |

## Inference Pipeline

<p align="center">
  <img src="assets/figure2_inference_pipeline.png" width="100%" alt="LLaDA-UI multimodal block-wise diffusion inference pipeline">
</p>

<p align="center"><em>GUI observations from web, mobile, and desktop environments are encoded at native resolution and combined with task and interaction-history tokens. The LLaDA2.0 decoder progressively denoises the model output into reasoning and executable actions.</em></p>

## Benchmark Results

<p align="center">
  <img src="assets/gui_agent_evaluation.png" width="100%" alt="LLaDA-UI GUI-agent benchmark results">
</p>

<p align="center"><em>GUI-agent evaluation reproduced from the technical report. See the code repository for evaluation details and the latest result artifacts.</em></p>

## Quick Start

### 1. Clone the inference code

```bash
git clone https://github.com/inclusionAI/LLaDA-UI.git
cd LLaDA-UI
```

### 2. Create an environment

The standalone Hugging Face inference path has been tested with Python 3.10,
PyTorch 2.5.1, Transformers 4.51.0, and FlashAttention 2.7.4.post1.

```bash
conda create -n llada-ui python=3.10 -y
conda activate llada-ui

pip install torch==2.5.1 torchvision \
  --index-url https://download.pytorch.org/whl/cu124
pip install transformers==4.51.0 Pillow numpy einops accelerate \
  sentencepiece protobuf safetensors
pip install ninja
pip install flash-attn==2.7.4.post1 --no-build-isolation --no-cache-dir
```

FlashAttention must be built against a compatible CUDA and PyTorch setup. The
checkpoint is about 32 GB in BF16; additional GPU memory is required for model
execution and visual tokens.

### 3. Run single-image grounding

The following command downloads the checkpoint from Hugging Face on first use
and returns the center point of the requested UI element:

```bash
CUDA_VISIBLE_DEVICES=0 IMAGE_MAX_PIXELS=12845056 \
python -u inference/inference_hf.py \
  --ckpt inclusionAI/LLaDA-UI \
  --image /path/to/screenshot.png \
  --prompt "click the search button" \
  --gen-length 32 \
  --steps 32 \
  --block-length 32
```

Example output:

```text
GENERATION: [742,186]
POINT(0~1): [0.742, 0.186]
```

The raw model coordinate is normalized to `[0, 999]`. The inference script
also prints the corresponding `[0, 1]` coordinate. To convert it to screen
pixels for an image of width `W` and height `H`, use `(x * W, y * H)` with the
`[0, 1]` values.

Large screenshots may also require increasing `--max-length` beyond its
default value of 8192.

## SGLang Serving

For agent evaluation or an OpenAI-compatible endpoint, follow the
[SGLang server guide](https://github.com/inclusionAI/LLaDA-UI/tree/main/inference/sglang_server).
The server uses LLaDA-UI-specific multimodal diffusion patches and cannot be
started from a stock SGLang installation alone.

After installing the documented serving environment, launch it with:

```bash
huggingface-cli download inclusionAI/LLaDA-UI \
  --local-dir /path/to/LLaDA-UI-checkpoint

CUDA_VISIBLE_DEVICES=0,1 SGLANG_DP_SIZE=2 \
bash serve_llada_ui.sh /path/to/LLaDA-UI-checkpoint
```

Then run one of the packaged requests:

```bash
export SGLANG_BASE_URL=http://127.0.0.1:30000/v1
export SGLANG_MODEL=LLaDA-UI

python3 inference/sglang_client.py --example mobile
python3 inference/sglang_client.py --example desktop
python3 inference/sglang_client.py --example web
```

## Output Format

For GUI navigation, a typical response contains tagged reasoning followed by
a structured action:

```text
<think>Reasoning grounded in the current GUI state.</think>
<action>Click(box=(x,y))</action>
```

Here, `x` and `y` are integer-normalized to `[0, 999]`. Despite the field name
`box`, point-based actions contain one interaction coordinate rather than a
rectangular bounding box.

For GUI grounding, the model directly returns `[x,y]`. It returns `[-1,-1]`
when the requested target is infeasible or unrelated to the screenshot.

## Citation

```bibtex
@misc{gu2026lladauibringingblockwisediffusion,
  title={LLaDA-UI: Bringing Block-wise Diffusion to Vision-Language GUI Agents},
  author={Zhangxuan Gu and Haoxing Chen and Qi Qin and Yi Xin and Kai Gan and Lin Liu and Long Cui and Xiaomei Wang and Beitong Zhou and Yunzhu Zhang and Zhengwen Zeng and Changlong Gao and Weizhi Chen and Rongchao Zhang and Haoyuan Wu and Shuheng Shen and Changhua Meng and Weiqiang Wang and Jianguo Li and Zhenzhong Lan},
  year={2026},
  eprint={2609.13287},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2609.13287},
}
```