|
|
--- |
|
|
license: mit |
|
|
language: |
|
|
- en |
|
|
base_model: |
|
|
- Qwen/Qwen2.5-VL-7B-Instruct |
|
|
tags: |
|
|
- arxiv:2508.15144 |
|
|
--- |
|
|
|
|
|
# GUI-Owl |
|
|
|
|
|
<div align="center"> |
|
|
<img src=https://youke1.picui.cn/s1/2025/08/18/68a2f82fef3d4.png width="40%"/> |
|
|
</div> |
|
|
|
|
|
GUI-Owl is a model series developed as part of the Mobile-Agent-V3 project. It achieves state-of-the-art performance across a range of GUI automation benchmarks, including ScreenSpot-V2, ScreenSpot-Pro, OSWorld-G, MMBench-GUI, Android Control, Android World, and OSWorld. Furthermore, it can be instantiated as various specialized agents within the Mobile-Agent-V3 multi-agent framework to accomplish more complex tasks. |
|
|
|
|
|
* **Paper**: [Paper Link](https://github.com/X-PLUG/MobileAgent/blob/main/Mobile-Agent-v3/assets/MobileAgentV3_Tech.pdf) |
|
|
* **GitHub Repository**: https://github.com/X-PLUG/MobileAgent |
|
|
* **Online Demo**: Comming soon |
|
|
|
|
|
This is a variant version of GUI-Owl that is specially RL-tuned for desktop environments. You can deploy and invoke the model using the same methods as GUI-Owl 7B. |
|
|
|
|
|
Limitation: This model is primarily intended for validating improvements and optimizations for environment-rl. Therefore, its performance in production scenarios cannot be guaranteed. |
|
|
|
|
|
|
|
|
## Usage |
|
|
|
|
|
Please refer to our cookbook. |
|
|
|
|
|
## Deploy |
|
|
|
|
|
We recommand deploy GUI-Owl-7B through vllm |
|
|
|
|
|
This script has been validated on an A100 with 96 GB of VRAM. |
|
|
```bash |
|
|
PIXEL_ARGS='{"min_pixels":3136,"max_pixels":10035200}' |
|
|
IMAGE_LIMIT_ARGS='image=2' |
|
|
MP_SIZE=1 |
|
|
MM_KWARGS=( |
|
|
--mm-processor-kwargs $PIXEL_ARGS |
|
|
--limit-mm-per-prompt $IMAGE_LIMIT_ARGS |
|
|
) |
|
|
|
|
|
vllm serve $CKPT \ |
|
|
--max-model-len 32768 ${MM_KWARGS[@]} \ |
|
|
--tensor-parallel-size $MP_SIZE \ |
|
|
--allowed-local-media-path '/' \ |
|
|
--port 4243 |
|
|
``` |
|
|
|
|
|
If you want GUI-Owl to recieve more than two images, you could increase `IMAGE_LIMIT_ARGS` and reduce `max_pixels`. |
|
|
|
|
|
For example: |
|
|
```bash |
|
|
PIXEL_ARGS='{"min_pixels":3136,"max_pixels":3211264}' |
|
|
IMAGE_LIMIT_ARGS='image=5' |
|
|
MP_SIZE=1 |
|
|
MM_KWARGS=( |
|
|
--mm-processor-kwargs $PIXEL_ARGS |
|
|
--limit-mm-per-prompt $IMAGE_LIMIT_ARGS |
|
|
) |
|
|
|
|
|
vllm serve $CKPT \ |
|
|
--max-model-len 32768 ${MM_KWARGS[@]} \ |
|
|
--tensor-parallel-size $MP_SIZE \ |
|
|
--allowed-local-media-path '/' \ |
|
|
--port 4243 |
|
|
``` |
|
|
|
|
|
## Citation |
|
|
If you find our paper and model useful in your research, feel free to give us a cite. |
|
|
``` |
|
|
@misc{ye2025mobileagentv3foundamentalagentsgui, |
|
|
title={Mobile-Agent-v3: Foundamental Agents for GUI Automation}, |
|
|
author={Jiabo Ye and Xi Zhang and Haiyang Xu and Haowei Liu and Junyang Wang and Zhaoqing Zhu and Ziwei Zheng and Feiyu Gao and Junjie Cao and Zhengxi Lu and Jitong Liao and Qi Zheng and Fei Huang and Jingren Zhou and Ming Yan}, |
|
|
year={2025}, |
|
|
eprint={2508.15144}, |
|
|
archivePrefix={arXiv}, |
|
|
primaryClass={cs.AI}, |
|
|
url={https://arxiv.org/abs/2508.15144}, |
|
|
} |
|
|
``` |
|
|
|