Image-Text-to-Text
PEFT
Safetensors
English
code
code-generation
coding-agent
agent
agentic
software-engineering
swe-bench
lora
qwen
qwen3
multimodal
vision
visual-coding
image-to-code
chart-to-code
render-verified
execution-verified
best-of-n
self-hostable
open-source
australia
Instructions to use SwarmDo/SwarmDo-A2 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use SwarmDo/SwarmDo-A2 with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen3.6-27B") model = PeftModel.from_pretrained(base_model, "SwarmDo/SwarmDo-A2") - Notebooks
- Google Colab
- Kaggle
File size: 10,522 Bytes
064d1a6 47a14b7 064d1a6 4650679 064d1a6 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 | ---
license: apache-2.0
language:
- en
base_model: Qwen/Qwen3.6-27B
base_model_relation: adapter
library_name: peft
pipeline_tag: image-text-to-text
tags:
- code
- code-generation
- coding-agent
- agent
- agentic
- software-engineering
- swe-bench
- lora
- peft
- qwen
- qwen3
- multimodal
- vision
- visual-coding
- image-to-code
- chart-to-code
- image-text-to-text
- render-verified
- execution-verified
- best-of-n
- self-hostable
- open-source
- australia
---
# SwarmDo-A2 β one model, both skills (code **and** vision)
**SwarmDo-A2 is an open, self-hostable, multimodal coding agent that fixes real software bugs *and* writes code from images β in a single model.** It combines SwarmDo-A1's execution-verified **coding** delta with a teacher-free, render-verified **visual-coding** delta, folded together into one LoRA adapter on `Qwen/Qwen3.6-27B`. Built in Australia; measured, not marketing β we publish the negatives and the confidence intervals.
> Public model hub & full results: **https://github.com/SwarmDo/models**
**Keywords:** open-source coding agent Β· agentic software engineering Β· SWE-bench-style patch generation Β· execution-verified best-of-N selection Β· image-to-code / chart-to-code / visual coding Β· render-verified self-improvement Β· multimodal vision LLM Β· Qwen3.6-27B LoRA / PEFT adapter Β· self-hostable on vLLM Β· Apache-2.0.
## What A2 is (and how it differs from A1)
[SwarmDo-A1](https://huggingface.co/SwarmDo/SwarmDo-A1) is the **coding** tier β a re-based open base plus a hardened agentic harness that verifies patches by **running the project's tests**. SwarmDo-**A2** takes that same coding delta and **combines it, on the same 27B base, with a separately trained visual-coding delta** β a *render-verified flywheel* that turns a chart / plot / UI screenshot into the code that reproduces it and checks itself by **rendering the output and comparing to the target image** (an objective, hard-to-hack reward). The two deltas are fused into one rank-48 adapter, so A2 is a **single checkpoint that carries both skills** rather than two models you swap between.
The adapter reaches the architecture's **gated-DeltaNet sequence-mixing layers** (`in_proj_*`, `out_proj`), not just the standard attention projections β which is why A2 must be **served merged** (see Usage).
## What's decision-grade (paired, held-out, execution-verified)
The headline claim for A2 is the **churn gate**: folding two independently trained skills into one model must not degrade *either* skill. We measured both, paired and held-out.
| Result | Number | Significance |
|---|---|---|
| **Visual moat preserved** β combined A2 vs the pure visual model (held-out chart render-reward) | Ξ β0.006, CI [β0.016, +0.004] | p = 0.63 (**not** worse) |
| Combined A2 still beats the raw base on visual coding | 14 wins : 5 | directional |
| **Coding preserved** β A2 vs A1 on held-out SWE tasks (**execution-grounded**, real `FAIL_TO_PASS`) | 18/40 vs 19/40 solved, Ξ β1 | McNemar p = 1.0 (39/40 agreement) |
| Visual flywheel **compounds** (teacher-free self-fuel β held-out lift) | +0.0192 render-reward | sign-test p = 0.0063 |
Inherited from the SwarmDo line (see A1): re-basing to Qwen3.6-27B β **+25 pp** (p = 0.013); base **+ execution-verified selection** solves **36 vs 24** (p = 0.0005); visual coding is image-driven (image-on 15 vs image-off 0, p = 6e-05).
These are **relative, held-out, execution-verified** comparisons on our own task sets β **not** an absolute public-leaderboard number we can't stand behind under paired testing.
> **Methodology note we're proud of:** our harness can *self-judge* a patch (an LLM reading its own diff). We found that signal is ~2Γ optimistic, so **every SWE number above is re-derived by actually executing the project's tests** against the generated patch β never by self-grading.
## Honest limitations (published, not buried)
- **As a single model, A2 is ~neutral vs the base on raw coding** β the same "churn wall" trade-off A1 documents. A2's value is (1) **carrying both skills in one model** with *no measured degradation of either*, and (2) the **system** around it (execution-verified best-of-N selection + render-verified visual coding).
- The **SWE churn arm is deliberately small** (40 paired tasks, 1 discordant pair): it rules out *gross* coding regression (39/40 verdict agreement, 40% byte-identical patches) but is underpowered for a *subtle* difference. The evidence for "coding preserved" is the near-identity, not the p-value.
- Python-centric agentic SWE + chart/plot/UI visual coding. Like all coding agents it can produce plausible-but-wrong patches β **always run the project's tests against its output.** Not a general chat assistant.
## Usage β serve **merged** (important)
A2's adapter targets the base's gated-DeltaNet modules, which **JIT-hang when served as a vLLM LoRA**. Serve A2 as **merged weights** (native fast path). Merge with the VL-aware helper (it preserves the vision tower and emits the image-processor config so vLLM can serve it):
```bash
# 1) merge the A2 adapter into the base, keeping the vision tower
python merge_lora_vl.py --base Qwen/Qwen3.6-27B --lora ./ --out ./swarmdo-a2-merged
# 2) serve the merged model (multimodal + tool-calling)
vllm serve ./swarmdo-a2-merged --served-model-name swarmdo-a2 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 \
--enforce-eager --trust-remote-code
```
`merge_lora_vl.py` is included in this repo. **For one-step serving with no merge, use the pre-merged weights β [SwarmDo/SwarmDo-A2-merged](https://huggingface.co/SwarmDo/SwarmDo-A2-merged)**:
```bash
vllm serve SwarmDo/SwarmDo-A2-merged --served-model-name swarmdo-a2 \
--enable-auto-tool-choice --tool-call-parser qwen3_xml --reasoning-parser qwen3 \
--enforce-eager --trust-remote-code
```
Or load with PEFT for local experimentation:
```python
from transformers import AutoModelForImageTextToText, AutoProcessor
from peft import PeftModel
base = AutoModelForImageTextToText.from_pretrained("Qwen/Qwen3.6-27B", torch_dtype="bfloat16", device_map="auto", trust_remote_code=True)
model = PeftModel.from_pretrained(base, "SwarmDo/SwarmDo-A2").merge_and_unload()
processor = AutoProcessor.from_pretrained("Qwen/Qwen3.6-27B", trust_remote_code=True)
```
**Serving notes:** use `--tool-call-parser qwen3_xml` (not `hermes` β it silently drops tool calls) *with* `--enable-auto-tool-choice`; serve in eager mode if you hit hybrid-attention graph-capture issues; the linear-attention base keeps a fixed-size recurrent state, so long contexts cost prefill tokens, not growing KV cache. A **text-only GGUF is not offered for A2** β quantizing to GGUF drops the vision tower, which would make the visual half inert. For code-only local runs, use [SwarmDo-A1-GGUF](https://huggingface.co/SwarmDo/SwarmDo-A1-GGUF).
## Use cases
- **One model for both** automated bug-fixing *and* image-to-code β no adapter swapping.
- **Automated bug fixing** in Python repositories β generate a patch, verify it against the project's tests.
- **Image-to-code / chart-to-code / screenshot-to-code** β reproduce a chart, plot, or UI as working code, checked by rendering and comparing.
- **Agentic software engineering** β a coding agent that edits files, runs commands, and self-checks patches.
- **Self-hosted / private multimodal coding assistant** β runs entirely on your own hardware via vLLM; no data leaves your machine.
- **Best-of-N patch generation** β sample multiple candidate fixes and select the one that passes execution.
## License
Apache-2.0, consistent with the base model family.
## Citation
```bibtex
@software{swarmdo_a2_2026,
title = {SwarmDo-A2: One open model for execution-verified coding and render-verified visual coding},
author = {SwarmDo},
year = {2026},
url = {https://github.com/SwarmDo/models}
}
```
## FAQ
**What is SwarmDo-A2?** An open, self-hostable multimodal coding agent that both **fixes software bugs** and **writes code from images**, in a single model β the coding (A1) delta and a render-verified visual-coding delta fused into one LoRA on `Qwen/Qwen3.6-27B`.
**How is A2 different from A1?** A1 is the coding tier (adapter + execution-verified harness, with a text GGUF for easy local runs). A2 folds a *separately trained visual-coding skill* into the same model, so one checkpoint does code **and** vision β with measured proof that combining them degrades neither skill.
**Is it open source?** Yes β Apache-2.0, commercial use permitted. Weights, recipe, and full (including negative) results are public.
**Can I run it on my own hardware?** Yes β self-host with vLLM on the **merged** weights (see Usage: a one-line merge, then `vllm serve`). Full multimodal serving needs a GPU sized for a 27B; a driver new enough for the base's CUDA build is required.
**Why must it be served merged, not as a LoRA?** The adapter reaches the base's gated-DeltaNet layers, which JIT-hang under vLLM's LoRA path. Merged weights use the native fast path.
**Is there a GGUF?** Not for A2 β GGUF conversion drops the vision tower, making the visual half inert. Use A2 merged (vLLM) for code+vision, or the [A1 GGUF](https://huggingface.co/SwarmDo/SwarmDo-A1-GGUF) for code-only local runs.
**Does it really do image-to-code?** Yes β it writes code from images and we verify it by *rendering the output and comparing to the target image*, an objective signal, not eyeballing. The visual skill is a teacher-free flywheel that measurably compounds as we scale its self-generated fuel.
**Is it good at SWE-bench?** We report *relative, paired, held-out, execution-verified* results (and re-derive every SWE number by running real tests, not self-grading), not a single absolute leaderboard number we can't stand behind.
**What languages?** Python-centric agentic software engineering, plus Python visual coding (charts/plots/UI). It is a coding agent, not a general chat assistant.
**Where are the full results?** Public hub: https://github.com/SwarmDo/models
---
*Topics: open source multimodal coding model Β· one model for code and vision Β· self-hosted AI coding assistant Β· agentic SWE-bench model Β· image to code / chart to code / screenshot to code Β· render-verified visual coding Β· execution-verified code generation Β· Qwen3.6-27B LoRA Β· run a coding LLM locally Β· Apache-2.0 code model.*
|