drunksu commited on
Commit
befdd31
·
verified ·
1 Parent(s): 41c488c

Update model card: execution wording, training/eval split, citation

Browse files
Files changed (1) hide show
  1. README.md +15 -8
README.md CHANGED
@@ -14,14 +14,15 @@ tags:
14
  - reinforcement-learning
15
  ---
16
 
17
- # GUISwiper
18
 
19
- GUISwiper is an RL-aligned GUI agent model for **human-like swipe synthesis**, introduced in the paper
20
  [SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis](https://arxiv.org/abs/2601.18305)
21
  (ACM MM 2026 Oral).
22
 
23
  This repository hosts the **final RL-aligned 3B checkpoint** (bfloat16), fine-tuned from
24
- [`Qwen/Qwen2.5-VL-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct).
 
25
 
26
  ## Model Details
27
 
@@ -29,8 +30,9 @@ This repository hosts the **final RL-aligned 3B checkpoint** (bfloat16), fine-tu
29
  |-----------------|----------------------------------------------|
30
  | Base model | Qwen/Qwen2.5-VL-3B-Instruct |
31
  | Parameters | 3B (bfloat16, ~7 GB) |
32
- | Training | SFT + RL alignment on SwipeBench |
33
- | Input | GUI screenshots + instruction |
 
34
  | Output | Human-like swipe action (trajectory) |
35
  | Hardware | NVIDIA GPUs (see paper for details) |
36
 
@@ -54,7 +56,7 @@ processor = AutoProcessor.from_pretrained(repo_id)
54
  image = load_your_gui_screenshot() # PIL.Image
55
  messages = [{"role": "user", "content": [
56
  {"type": "image", "image": image},
57
- {"type": "text", "text": "Swipe to learn more."},
58
  ]}]
59
  text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
60
  inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
@@ -80,17 +82,22 @@ path = hf_hub_download("drunksu/GUISwiper", "model-00001-of-00002.safetensors")
80
  title = {SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis},
81
  author = {SwipeGen Team},
82
  journal = {arXiv preprint arXiv:2601.18305},
83
- year = {2026}
 
84
  }
85
  ```
86
 
 
 
 
87
  ## Links
88
 
89
  - Paper: https://arxiv.org/abs/2601.18305
90
  - Project / code: https://github.com/TSKGHS17/SwipeGen
 
91
 
92
  ## License & Disclaimer
93
 
94
  The model weights are released under Apache-2.0, consistent with the base model
95
  `Qwen2.5-VL-3B-Instruct`. Users should comply with the original license terms of Qwen2.5-VL
96
- and use the model responsibly; outputs are generated by AI and may contain errors.
 
14
  - reinforcement-learning
15
  ---
16
 
17
+ # GUISwiper 3B (RL)
18
 
19
+ GUISwiper is an RL-aligned GUI agent model for **human-like swipe execution**, introduced in the paper
20
  [SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis](https://arxiv.org/abs/2601.18305)
21
  (ACM MM 2026 Oral).
22
 
23
  This repository hosts the **final RL-aligned 3B checkpoint** (bfloat16), fine-tuned from
24
+ [`Qwen/Qwen2.5-VL-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct) and
25
+ evaluated on [SwipeBench](https://huggingface.co/datasets/TSKGHS17/SwipeBench).
26
 
27
  ## Model Details
28
 
 
30
  |-----------------|----------------------------------------------|
31
  | Base model | Qwen/Qwen2.5-VL-3B-Instruct |
32
  | Parameters | 3B (bfloat16, ~7 GB) |
33
+ | Training | SFT + RL alignment (see paper for details) |
34
+ | Evaluation | SwipeBench |
35
+ | Input | GUI screenshots / screen videos + instruction |
36
  | Output | Human-like swipe action (trajectory) |
37
  | Hardware | NVIDIA GPUs (see paper for details) |
38
 
 
56
  image = load_your_gui_screenshot() # PIL.Image
57
  messages = [{"role": "user", "content": [
58
  {"type": "image", "image": image},
59
+ {"type": "text", "text": "Describe the swipe to perform here."},
60
  ]}]
61
  text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
62
  inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
 
82
  title = {SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis},
83
  author = {SwipeGen Team},
84
  journal = {arXiv preprint arXiv:2601.18305},
85
+ year = {2026},
86
+ note = {Code and models: \url{https://github.com/TSKGHS17/SwipeGen}}
87
  }
88
  ```
89
 
90
+ If you use GUISwiper, please also reference the official repository:
91
+ <https://github.com/TSKGHS17/SwipeGen>.
92
+
93
  ## Links
94
 
95
  - Paper: https://arxiv.org/abs/2601.18305
96
  - Project / code: https://github.com/TSKGHS17/SwipeGen
97
+ - Dataset (SwipeBench): https://huggingface.co/datasets/TSKGHS17/SwipeBench
98
 
99
  ## License & Disclaimer
100
 
101
  The model weights are released under Apache-2.0, consistent with the base model
102
  `Qwen2.5-VL-3B-Instruct`. Users should comply with the original license terms of Qwen2.5-VL
103
+ and use the model responsibly; outputs are generated by AI and may contain errors.