Experiencency's picture
Tighten base model address spacing
c92fa07 verified
|
Raw
History Blame Contribute Delete
1.88 kB
---
license: apache-2.0
base_model: GELab-Zero-4B-preview
library_name: transformers
pipeline_tag: image-text-to-text
language:
- en
- zh
tags:
- gui-agent
- mobile-agent
- vision-language
- qwen3-vl
- lora
---
> - 🔧 This model is part of [**Sico**](https://github.com/microsoft/Sico) — an open-source platform for building and evolving Digital Workers, where AI agents and their human operators co-evolve through real work.
> - ⭐ [**Star the Sico repository**](https://github.com/microsoft/Sico) to follow new evolved models and our model evolution pipeline — this GUI agent is the first public release, with more on the way.
> - 📄 Backed by our survey on [**agentic evolution and co-evolving human–AI systems**](https://www.microsoft.com/en-us/research/publication/agentic-evolution-from-self-improving-agents-to-co-evolving-human-ai-systems/).
# GELab-Zero-4B-preview-Sico-Evolution
A 4B GUI agent fine-tuned (LoRA) from the open-source **GELab-Zero-4B-preview** base model
on **Microsoft Edge** and **Copilot** UI trajectories. It is built with our **general-purpose
GUI model evolution pipeline** — an iterative mechanism that keeps lifting an agent's real
task success rate round after round, and transfers to any GUI app.
**Base model address:** https://huggingface.co/stepfun-ai/GELab-Zero-4B-preview
## Highlights
**From 39.8% to 82.9%:** Sico-Evolution achieves a dominant **82.9%
Task Success Rate**, a massive **+43.1% absolute surge** over the **39.8%** base-model
baseline.
**Outperforms Closed-Source SOTAs:** It edges out top proprietary giants like **gpt-5.4
(79.7%)**, **Claude-Opus-4.6 (81.3%)**, and **claude-opus-4.7 (82.1%)**.
**Vastly Exceeds Open-Source Models:** It crushes leading competitors including
**kimi-k2.6 (62.6%)** and **UI-Venus-1.5-30B (61.0%)**.
## Results
![Edge / Copilot Test Cases — TSR](tsr_edge_copilot.png)