drunksu commited on
Commit
41c488c
·
verified ·
1 Parent(s): 461b647

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +5 -9
README.md CHANGED
@@ -1,8 +1,6 @@
1
  ---
2
  license: apache-2.0
3
  base_model: Qwen/Qwen2.5-VL-3B-Instruct
4
- datasets:
5
- - TSKGHS17/SwipeBench
6
  language:
7
  - en
8
  - zh
@@ -16,15 +14,14 @@ tags:
16
  - reinforcement-learning
17
  ---
18
 
19
- # GUISwiper 3B (RL)
20
 
21
  GUISwiper is an RL-aligned GUI agent model for **human-like swipe synthesis**, introduced in the paper
22
  [SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis](https://arxiv.org/abs/2601.18305)
23
  (ACM MM 2026 Oral).
24
 
25
  This repository hosts the **final RL-aligned 3B checkpoint** (bfloat16), fine-tuned from
26
- [`Qwen/Qwen2.5-VL-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct) on
27
- [SwipeBench](https://huggingface.co/datasets/TSKGHS17/SwipeBench).
28
 
29
  ## Model Details
30
 
@@ -33,7 +30,7 @@ This repository hosts the **final RL-aligned 3B checkpoint** (bfloat16), fine-tu
33
  | Base model | Qwen/Qwen2.5-VL-3B-Instruct |
34
  | Parameters | 3B (bfloat16, ~7 GB) |
35
  | Training | SFT + RL alignment on SwipeBench |
36
- | Input | GUI screenshots / screen videos + instruction |
37
  | Output | Human-like swipe action (trajectory) |
38
  | Hardware | NVIDIA GPUs (see paper for details) |
39
 
@@ -57,7 +54,7 @@ processor = AutoProcessor.from_pretrained(repo_id)
57
  image = load_your_gui_screenshot() # PIL.Image
58
  messages = [{"role": "user", "content": [
59
  {"type": "image", "image": image},
60
- {"type": "text", "text": "Describe the swipe to perform here."},
61
  ]}]
62
  text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
63
  inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
@@ -91,10 +88,9 @@ path = hf_hub_download("drunksu/GUISwiper", "model-00001-of-00002.safetensors")
91
 
92
  - Paper: https://arxiv.org/abs/2601.18305
93
  - Project / code: https://github.com/TSKGHS17/SwipeGen
94
- - Dataset (SwipeBench): https://huggingface.co/datasets/TSKGHS17/SwipeBench
95
 
96
  ## License & Disclaimer
97
 
98
  The model weights are released under Apache-2.0, consistent with the base model
99
  `Qwen2.5-VL-3B-Instruct`. Users should comply with the original license terms of Qwen2.5-VL
100
- and use the model responsibly; outputs are generated by AI and may contain errors.
 
1
  ---
2
  license: apache-2.0
3
  base_model: Qwen/Qwen2.5-VL-3B-Instruct
 
 
4
  language:
5
  - en
6
  - zh
 
14
  - reinforcement-learning
15
  ---
16
 
17
+ # GUISwiper
18
 
19
  GUISwiper is an RL-aligned GUI agent model for **human-like swipe synthesis**, introduced in the paper
20
  [SwipeGen: Bridging the Execution Gap in GUI Agents via Human-like Swipe Synthesis](https://arxiv.org/abs/2601.18305)
21
  (ACM MM 2026 Oral).
22
 
23
  This repository hosts the **final RL-aligned 3B checkpoint** (bfloat16), fine-tuned from
24
+ [`Qwen/Qwen2.5-VL-3B-Instruct`](https://huggingface.co/Qwen/Qwen2.5-VL-3B-Instruct).
 
25
 
26
  ## Model Details
27
 
 
30
  | Base model | Qwen/Qwen2.5-VL-3B-Instruct |
31
  | Parameters | 3B (bfloat16, ~7 GB) |
32
  | Training | SFT + RL alignment on SwipeBench |
33
+ | Input | GUI screenshots + instruction |
34
  | Output | Human-like swipe action (trajectory) |
35
  | Hardware | NVIDIA GPUs (see paper for details) |
36
 
 
54
  image = load_your_gui_screenshot() # PIL.Image
55
  messages = [{"role": "user", "content": [
56
  {"type": "image", "image": image},
57
+ {"type": "text", "text": "Swipe to learn more."},
58
  ]}]
59
  text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
60
  inputs = processor(text=[text], images=[image], return_tensors="pt").to(model.device)
 
88
 
89
  - Paper: https://arxiv.org/abs/2601.18305
90
  - Project / code: https://github.com/TSKGHS17/SwipeGen
 
91
 
92
  ## License & Disclaimer
93
 
94
  The model weights are released under Apache-2.0, consistent with the base model
95
  `Qwen2.5-VL-3B-Instruct`. Users should comply with the original license terms of Qwen2.5-VL
96
+ and use the model responsibly; outputs are generated by AI and may contain errors.