File size: 14,118 Bytes
b16888d
65cfdff
 
374dfa0
65cfdff
 
 
 
b16888d
65cfdff
c96b5c2
 
 
 
65cfdff
 
 
 
374dfa0
6a54615
 
65cfdff
 
 
 
 
 
 
 
374dfa0
65cfdff
 
 
374dfa0
65cfdff
 
 
c96b5c2
 
 
 
65cfdff
 
 
 
 
374dfa0
65cfdff
 
 
 
 
374dfa0
 
65cfdff
 
c96b5c2
 
 
 
374dfa0
c96b5c2
 
 
 
 
65cfdff
 
 
 
374dfa0
 
65cfdff
 
 
 
 
 
 
 
 
 
 
 
374dfa0
 
 
 
 
 
 
 
6a54615
374dfa0
 
 
 
6a54615
 
 
374dfa0
65cfdff
 
 
258cb94
65cfdff
 
 
374dfa0
 
 
 
258cb94
 
 
 
 
374dfa0
 
 
 
65cfdff
 
 
374dfa0
65cfdff
374dfa0
 
258cb94
 
 
 
 
 
374dfa0
 
 
 
65cfdff
 
 
 
 
374dfa0
 
258cb94
 
 
 
 
374dfa0
 
 
 
65cfdff
 
 
 
 
 
 
258cb94
 
 
 
 
374dfa0
 
 
 
65cfdff
 
 
374dfa0
65cfdff
374dfa0
 
258cb94
 
 
 
374dfa0
 
 
 
65cfdff
 
 
374dfa0
65cfdff
374dfa0
 
 
 
258cb94
 
 
 
 
 
374dfa0
 
65cfdff
 
 
374dfa0
65cfdff
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
---
pipeline_tag: image-text-to-text
library_name: transformers
base_model: Qwen/Qwen3.5-9B
tags:
  - multimodal
  - gui
  - agent
---

<div align="center">
  <img src="assets/ui-venus-logo.png" width="140" alt="UI-Venus logo">
</div>

# UI-Venus-2-9B

<div align="center">

[![Technical Report](https://img.shields.io/badge/Technical%20Report-Coming%20Soon-lightgrey.svg)](#citation)
[![Website](https://img.shields.io/badge/Website-Project%20Page-blue.svg)](https://ui-venus.github.io/UI-Venus-2/)
[![Code](https://img.shields.io/badge/Code-GitHub-black.svg)](https://github.com/inclusionAI/UI-Venus)
[![Hugging Face](https://img.shields.io/badge/🤗%20Hugging%20Face-Collection-yellow.svg)](https://huggingface.co/collections/inclusionAI/ui-venus)

</div>

**UI-Venus-2** is a general-purpose foundation GUI agent designed to operate across **mobile applications, web platforms, and desktop operating systems** through a unified closed-loop reasoning–action framework: the agent observes the current interface, reasons about the task state, executes an action, and incorporates environmental feedback into its next decision.

To bridge the gap from benchmark-oriented models toward dependable real-world applications, UI-Venus-2 jointly scales three critical dimensions:

- **Environments** — expanding the executable mobile pool to 170+ multilingual apps (100+ Chinese and 70+ English), alongside a curated web pool of 4,000+ domains across 19 categories and native desktop operating systems;
- **Tasks** — a deep-research pipeline that grounds generated instructions in the actual functionality of each application, improving the accuracy, validity, and executability of task queries;
- **Verification** — trace-level and sample-level evaluators based on task-relevant **visual keypoints** and **multi-model voting**, providing reliable RL reward signals that are robust to reward hacking.

UI-Venus-2 is evaluated for consequential-action safety on OSHarm and OSBlind. UI-Venus-2-9B achieves an Attack Success Rate (ASR, lower is better) of **11.3% on OSHarm** and **48.8% on OSBlind**, compared with 25.3% and 79.4%, respectively, for its Qwen3.5-9B base model.

This repository releases the full-parameter weights of **UI-Venus-2-9B**, initialized from [Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B), which achieves near state-of-the-art performance among models of comparable scale across GUI grounding, mobile, web, computer-use, and CAPTCHA benchmarks.

<div align="center">
  <img src="assets/performance_overview.png" width="100%" alt="Performance overview of UI-Venus-2 across GUI agent benchmarks">
</div>

## Highlights

- **Scaled multilingual mobile-use environments with reliable task generation.** A substantially expanded executable mobile pool covering Chinese and English app ecosystems, paired with a deep-research-driven query-generation strategy grounded in real application functionality.
- **Computer-use capability built from the ground up.** Dedicated desktop operating-system capabilities constructed from scratch through computer-use data collection and task-specific training, extending the UI-Venus family to mobile, web, and OS interaction in one unified end-to-end agent.
- **Keypoint-grounded verification.** Task completion is judged on task-relevant visual keypoints rather than a coarse holistic look at the final screen, with multi-model voting aggregating heterogeneous judges — reducing single-judge bias and making the reward signal robust to reward hacking.
- **Multi-teacher on-policy distillation.** Domain-specialized teachers are consolidated with structured, action-aware supervision that concentrates the distillation signal on the executable behaviors governing environment interaction.

## Training Pipeline

UI-Venus-2 is trained on a deliberately complementary mixture of five task families — **Grounding, CAPTCHA, Mobile, Web, and Computer** — through three stages:

1. **Multimodal Mid-Training** — large-scale trajectory-based training over simulated mobile, web, and OS environments, with navigation data as the dominant component and human–discriminator collaborative verification filtering invalid or ambiguous interactions.
2. **Offline RL** — independent step-level optimization across Grounding, CAPTCHA, Mobile, Web, and Computer tasks develops domain-specialized policies.
3. **Multi-teacher On-policy Distillation** — domain-specialized teachers are consolidated into a single unified policy that preserves the broad multimodal reasoning of the base model while composing spatial grounding, verified interaction, and long-horizon navigation.

<div align="center">
  <img src="assets/training_pipeline.png" width="90%" alt="The three-stage training pipeline of UI-Venus-2">
</div>

The training data is produced by a deep-research-driven task-generation pipeline spanning mobile, browser, and computer environments. The active report describes a curated web pool of 4,000+ domains across 19 categories, seeded with 45,000 tasks selected from InSTA-150k-v3; synthetic grounding and CAPTCHA records are exported for task-specific SFT and RL.

<div align="center">
  <img src="assets/system_overview.png" width="100%" alt="System overview: task generation pipeline and trajectory collection across mobile, browser, and computer environments">
</div>

## Quick Start

### Serving with vLLM

The following is the intended serving configuration. It was not live-canary-validated as part of this model-card update, so pin and verify a vLLM version appropriate for Qwen3.5 before deployment.

```bash
pip install -U vllm

vllm serve inclusionAI/UI-Venus-2-9B \
    --port 8000 \
    --tensor-parallel-size 1 \
    --max-model-len 262144 \
    --reasoning-parser qwen3
```

The server exposes an OpenAI-compatible API at `http://localhost:8000/v1`.

### Inference Configuration

The evaluation settings depend on the task family:

- **General agentic tasks:** use temperature `1.0`, enable reasoning, and retain the complete reasoning history in the multi-step context.
- **GUI grounding:** disable reasoning and use temperature `0` for direct point localization.
- **CAPTCHA:** use the CAPTCHA-specific prompt and a parser/executor that supports multiple actions in one model turn.

Starting the vLLM server alone does not provide the complete closed-loop GUI agent. Reference prompts, action parsers, and runnable examples for Mobile, Computer, Browser, Grounding, and CAPTCHA are available in the [UI-Venus code repository](https://github.com/inclusionAI/UI-Venus). Environment coverage differs by domain: the included Mobile and Computer examples support multi-turn inference over prerecorded screenshots, while real-device ADB automation, browser interaction, and benchmark-specific setup are documented separately in the repository.

## Release Status

- **Model weights:** included in this repository.
- **Code:** available in the [UI-Venus GitHub repository](https://github.com/inclusionAI/UI-Venus).
- **Project page:** available at [ui-venus.github.io/UI-Venus-2](https://ui-venus.github.io/UI-Venus-2/).
- **Technical report and UI-Venus-2 citation:** will be added when the report is publicly released.
- **License:** the model-weight license is pending final confirmation and will be added before public release. The Apache-2.0 declaration has intentionally not been carried over because the current upstream materials contain conflicting license statements.

## Results

The tables below compare the released 9B checkpoint and the 27B family checkpoint with a representative subset of general-purpose and GUI-specialized baselines from the technical report. `–` denotes an unreported result, and `*` denotes a baseline evaluated or reproduced by the UI-Venus authors under the stated protocol. Full comparisons, source citations, and metric definitions are provided in the technical report. Live-environment results may vary with evaluation date and environment state.

### Mobile Use

Reported score (%) across interactive mobile-agent benchmarks, including controlled simulations and real-application environments.

| Model | MobileGym | VenusBench-Mobile | AndroidWorld | MobileWorld | KnowUBench | MemGUI |
|---|:---:|:---:|:---:|:---:|:---:|:---:|
| Qwen3.5-9B | 9.0* | 15.3* | 57.8 | 18.0 (18.0)* | 33.3 | 6.2* |
| Qwen3.6-27B | 24.6* | 28.0* | 70.3 | 36.8 (41.9)* | – | 25.7* |
| Seed-2.0-Pro | 52.0 | 20.1* | – | 63.2 | 51.6 | 65.6* |
| UI-Venus-1.5-30B-A3B | 21.5* | 21.5 | 77.6 | 17.1 | – | 10.9* |
| Qwen-UI-Agent-27B | – | – | – | 82.1 (85.5) | – | – |
| **UI-Venus-2-9B (released)** | 52.7 | 46.5 | 80.2 | 65.8 (75.2) | 56.5 | 62.6 |
| UI-Venus-2-27B | 60.5 | 48.7 | 84.0 | 76.1 (82.9) | 59.7 | 70.3 |

MobileWorld reports GUI-only success rate on 117 tasks: the primary value uses a 50-step budget and the value in parentheses uses 100 steps. VenusBench-Mobile uses its 149-task primary pool. MemGUI reports Main Results pass@1.

### Computer Use

Reported score (%) on open-ended tasks in desktop environments; benchmark-specific definitions follow the table.

| Model | OSWorld-Verified | OSWorld 2.0 Binary Acc. | OSWorld 2.0 Partial Score | DeskCraft |
|---|:---:|:---:|:---:|:---:|
| Qwen3.5-9B | 41.8 | 0.0* | 2.5* | 14.6* |
| Qwen3.6-27B | 62.0 | 0.0* | 3.8* | 28.7* |
| Kimi-K2.6 | 73.1 | 1.9 | 7.1 | 41.4* |
| GPT-5.5 | 78.7 | 13.0 | 46.7 | – |
| Claude-Opus-4.8 | 83.4 | – | – | – |
| Qwen-UI-Agent-27B | 79.5 | – | – | – |
| **UI-Venus-2-9B (released)** | 70.8 | 0.0 | 7.5 | 48.0 |
| UI-Venus-2-27B | 80.5 | 2.8 | 13.2 | 55.5 |

OSWorld-Verified comparisons use the 361-task setting reported by their cited sources and may use model-specific action scaffolds, so they should be interpreted as benchmark-level references rather than controlled ablations. OSWorld 2.0 reports the official Binary Accuracy and Partial Score metrics over 108 tasks under a 150-step budget. DeskCraft is an author-evaluated aggregate over the 538-task union of the Standard and Interactive splits and differs from official split-level reporting.

### Web Navigation

Success rate (%) on end-to-end web navigation; Odysseys reports the average rubric score (%).

| Model | WebVoyager | Online-Mind2Web | REAL | Odysseys Avg. | Odysseys Perfect |
|---|:---:|:---:|:---:|:---:|:---:|
| Qwen3.5-9B | 46.9* | 27.3* | 18.2* | 42.6* | 13.5* |
| Qwen3.6-27B | 84.3* | 55.3* | 27.3* | 39.5* | 18.5* |
| Seed-2.0-Pro | 85.1* | 68.5* | 74.4* | 60.2* | 30.1* |
| Claude-Opus-4.6 | 88.0 | – | – | 68.9 | 44.5 |
| Fara1.5-27B | 89.3 | 72.3 | – | – | – |
| **UI-Venus-2-9B (released)** | 90.8 | 74.0 | 76.9 | 77.3 | 62.0 |
| UI-Venus-2-27B | 93.4 | 78.3 | 80.2 | 80.4 | 66.3 |

WebVoyager uses the refreshed 595-task split; live-site states may vary by evaluation date. Odysseys evaluates 200 live-web tasks with `gemini-3.1-flash-lite-preview` as the rubric judge and reports both the average fraction of satisfied rubric items and the percentage of tasks for which every rubric item is satisfied.

### GUI Grounding

Accuracy (%) on element localization across high-resolution professional software, mobile / web / desktop interfaces, and reasoning-heavy instructions.

| Model | ScreenSpot-Pro | VenusBench-GD | OSWorld-G-R | UI-Vision |
|---|:---:|:---:|:---:|:---:|
| Qwen3.6-27B | 65.2* | 67.7* | 76.9* | 58.3* |
| Qwen 3.7 Plus | 68.9 | 75.2* | 78.2 | 68.0 |
| Seed 2.1 Pro | 65.3 | 73.9* | 78.0 | 62.0 |
| UI-Venus-1.5-30B-A3B | 69.6 | 75.0 | 76.4 | 54.7 |
| Qwen-UI-Agent-27B | 76.6 | – | 78.5 | 70.0 |
| **UI-Venus-2-9B (released)** | 73.0 | 77.1 | 78.5 | 53.2 |
| UI-Venus-2-27B | 74.1 | 80.1 | 79.1 | 66.9 |

VenusBench-GD reports English-instruction micro-average point-in-box accuracy.

### CAPTCHA Solving

Pass@1 (%) across five CAPTCHA benchmarks. VenusBench-CAPTCHA spans eight real-world interaction types—including slider puzzles, image rotation, text and icon click, OCR entry, and visual reasoning—and requires the full multi-action sequence in a single turn.

| Model | VenusBench-CAPTCHA | MCA-Bench | Spatial-CAPTCHA-Bench | NextGen-CAPTCHAs | Open CaptchaWorld |
|---|:---:|:---:|:---:|:---:|:---:|
| Qwen3.5-9B | 28.3 | 30.4 | 4.9 | 2.8 | 36.4 |
| Qwen3.6-27B | 53.0 | 51.7 | 31.0 | 14.1 | 47.7 |
| Seed-2.0-Pro | 47.9 | 36.5 | 43.8 | 20.4 | 55.6 |
| Kimi-K2.6 | 39.7 | 38.7 | 24.8 | 7.2 | 47.8 |
| **UI-Venus-2-9B (released)** | 78.1 | 75.7 | 42.8 | 47.6 | 50.7 |
| UI-Venus-2-27B | 79.9 | 79.6 | 48.6 | 54.5 | 56.3 |

All values are Pass@1. VenusBench-CAPTCHA reports the micro-average over all 219 fixed examples. The public-benchmark evaluation uses 1,000 sampled MCA-Bench examples, 15 retained NextGen-CAPTCHAs task types, and 16 retained Open CaptchaWorld task types.

### Safety

OSHarm covers explicit safety threats including deliberate misuse, prompt injection, and model misbehavior. OSBlind tests benign-looking instructions whose execution context can produce harmful outcomes. Both report Attack Success Rate (ASR, lower is better).

| Model | OSHarm ASR (%) ↓ | OSBlind ASR (%) ↓ |
|---|:---:|:---:|
| Qwen3.5-9B | 25.3 | 79.4 |
| Qwen3.5-27B | 18.0 | 89.3 |
| Kimi-K2.6 | 32.0 | 93.6 |
| EvoCUA-8B | 39.3 | 85.3 |
| UI-TARS-1.5 | 36.0 | 83.3 |
| ScaleCUA | 25.3 | 84.7 |
| **UI-Venus-2-9B (released)** | **11.3** | 48.8 |
| UI-Venus-2-27B | 15.3 | **47.9** |

These results do not make harmful-action risk negligible. Deployments should keep human confirmation and application-level safeguards around consequential operations, and CAPTCHA capabilities should only be used in authorized environments.

## Citation

The UI-Venus-2 citation will be added when the technical report is public. For prior UI-Venus releases, please use:

```bibtex
@article{uivenus15,
  title   = {UI-Venus-1.5 Technical Report},
  author  = {{Venus Team, Ant Group}},
  journal = {arXiv preprint arXiv:2602.09082},
  year    = {2026}
}

@article{uivenus2025,
  title   = {UI-Venus Technical Report: Building High-performance UI Agents with RFT},
  author  = {{Venus Team, Ant Group}},
  journal = {arXiv preprint arXiv:2508.10833},
  year    = {2025}
}
```