Spaces:
Running on Zero
Running on Zero
| title: Fara1.5-9B Computer Use Agent | |
| emoji: 🖱️ | |
| colorFrom: green | |
| colorTo: gray | |
| sdk: gradio | |
| sdk_version: 6.15.1 | |
| app_file: app.py | |
| short_description: Visual grounding demo for Fara1.5-9B computer use agent | |
| python_version: "3.12" | |
| startup_duration_timeout: 30m | |
| # Fara1.5-9B Computer Use Agent — Visual Grounding Demo | |
| This Space demonstrates [Fara1.5-9B](https://huggingface.co/microsoft/Fara1.5-9B), | |
| a 9B multimodal **computer use agent** by Microsoft Research AI Frontiers. | |
| ## How it works | |
| 1. Upload a screenshot of a web page | |
| 2. Describe the task the agent should perform (e.g. "Click the search box and type 'wireless headphones'") | |
| 3. Fara1.5-9B analyzes the screenshot and predicts the next grounded action | |
| 4. The predicted action (click coordinate, text to type, scroll, etc.) is visualized on the image with a marker | |
| ## Technical details | |
| - **Model**: microsoft/Fara1.5-9B (9B parameters, bf16, ~18GB VRAM) | |
| - **Base**: Qwen3.5-9B, fine-tuned with FaraGen1.5 synthetic data | |
| - **Coordinate system**: Normalized 0–1000 per axis, mapped onto a 1440×900 viewport | |
| - **Output format**: Chain-of-thought reasoning followed by an XML-tagged `computer_use` tool call | |
| - **System prompt**: The verbatim system prompt from the model card | |
| The model predicts pixel-level click coordinates in a normalized 0–1000 space, | |
| which are rescaled to the actual screenshot dimensions for visualization. | |