Instructions to use Pengyu965/Back2Struct-Image2SVG-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Pengyu965/Back2Struct-Image2SVG-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Pengyu965/Back2Struct-Image2SVG-7B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Pengyu965/Back2Struct-Image2SVG-7B") model = AutoModelForMultimodalLM.from_pretrained("Pengyu965/Back2Struct-Image2SVG-7B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Pengyu965/Back2Struct-Image2SVG-7B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Pengyu965/Back2Struct-Image2SVG-7B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pengyu965/Back2Struct-Image2SVG-7B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Pengyu965/Back2Struct-Image2SVG-7B
- SGLang
How to use Pengyu965/Back2Struct-Image2SVG-7B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Pengyu965/Back2Struct-Image2SVG-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pengyu965/Back2Struct-Image2SVG-7B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Pengyu965/Back2Struct-Image2SVG-7B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Pengyu965/Back2Struct-Image2SVG-7B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Pengyu965/Back2Struct-Image2SVG-7B with Docker Model Runner:
docker model run hf.co/Pengyu965/Back2Struct-Image2SVG-7B
Back2Struct: Making Structured Images Editable Again
Recovering editable, object-level SVG / XML code from images of structured graphics.
- 📄 Paper: arXiv:2609.37016 — Back2Struct: Making Structured Images Editable Again
- 🔗 Project page: https://pengyu965.github.io/Back2Struct.github.io/
- 🤖 Model:
Pengyu965/Back2Struct-Image2SVG-7B - 📚 Dataset:
Pengyu965/StructHub
Back2Struct "makes structured images editable again" by directly recovering vector
graphics code (SVG / XML) from image representations. Given an image of a structured
graphic — a diagram, chart, flowchart, architecture figure, UML diagram, or schema —
it predicts semantically object-level SVG/XML that explicitly encodes text, shapes,
topology, and layout (not low-level pixel tracing), so the code imports straight into
tools like PowerPoint to edit, restyle, and reuse while preserving structure. Training
is supervised fine-tuning from Qwen2.5-VL-7B-Instruct, followed by GRPO with a
composite reward (syntactic validity, length fidelity, perceptual quality).
This project ships two artifacts — a model and a dataset — documented together in this shared card.
🤖 Model — Back2Struct-Image2SVG-7B
| Base | Qwen/Qwen2.5-VL-7B-Instruct (7B) |
| Stage 1 — SFT | Supervised fine-tuning on StructHub image→SVG pairs |
| Stage 2 — GRPO | Group-Relative Policy Optimization (TRL) |
| Reward | Composite, render-gated: syntactic validity (must render) × perceptual fidelity (baseline-subtracted DINOv2 similarity between the rendered prediction and the target), encouraging compilability, length consistency, and structural/semantic faithfulness |
| Max output | 8,192 tokens · LoRA (r=64, α=128), merged into the released weights |
📚 Dataset — StructHub
Structured images paired with their editable SVG/XML source code.
| Split | Description | # examples |
|---|---|---|
train |
Full training set (all sources, all token lengths), benchmark held out | 86,102 |
benchmark |
Held-out evaluation benchmark (starvector 569 + crawled 334 + nn_diagram 97) | 1,000 |
Fields: id (string), source (string), token_count (int32), image (embedded PNG),
svg (ground-truth SVG/XML). The benchmark is stratified by SVG length into easy
(≤2,048), medium (2,048–4,096), and hard (>4,096) tiers (333/333/334).
Results
StructHub benchmark, render-gated (a non-rendering prediction scores 0 on DINO/SSIM/GPT, 1 on LPIPS). SR = render success (%); GPT = 0–100 GPTScore.
| Model | SR ↑ | DINO ↑ | LPIPS ↓ | SSIM ↑ | GPT ↑ |
|---|---|---|---|---|---|
| Overall | |||||
| Qwen3-VL-30B-A3B | 18.8 | 0.1566 | 0.9067 | 0.1201 | 13.29 |
| StarVector-8B | 31.4 | 0.2694 | 0.7936 | 0.2170 | 21.05 |
| Qwen2.5-VL-7B (base) | 52.3 | 0.3527 | 0.7834 | 0.3276 | 27.56 |
| Qwen2.5-VL-7B (SFT) | 42.8 | 0.3715 | 0.7403 | 0.2891 | 33.49 |
| Back2Struct (RL) | 71.5 | 0.5622 | 0.6473 | 0.4966 | 51.85 |
| Easy | |||||
| Qwen3-VL-30B-A3B | 26.5 | 0.2231 | 0.8659 | 0.1702 | 19.51 |
| StarVector-8B | 46.1 | 0.3923 | 0.7023 | 0.3123 | 31.90 |
| Qwen2.5-VL-7B (base) | 69.3 | 0.4776 | 0.7123 | 0.4196 | 39.66 |
| Qwen2.5-VL-7B (SFT) | 56.0 | 0.4866 | 0.6630 | 0.3789 | 44.58 |
| Back2Struct (RL) | 84.0 | 0.6384 | 0.5956 | 0.5796 | 62.63 |
| Medium | |||||
| Qwen3-VL-30B-A3B | 19.6 | 0.1631 | 0.9004 | 0.1288 | 13.38 |
| StarVector-8B | 37.0 | 0.3232 | 0.7533 | 0.2583 | 25.27 |
| Qwen2.5-VL-7B (base) | 51.8 | 0.3546 | 0.7816 | 0.3306 | 26.88 |
| Qwen2.5-VL-7B (SFT) | 45.5 | 0.3985 | 0.7161 | 0.3091 | 36.50 |
| Back2Struct (RL) | 75.6 | 0.6088 | 0.6182 | 0.5275 | 55.14 |
| Hard | |||||
| Qwen3-VL-30B-A3B | 10.2 | 0.0836 | 0.9536 | 0.0614 | 7.01 |
| StarVector-8B | 11.1 | 0.0932 | 0.9249 | 0.0808 | 6.02 |
| Qwen2.5-VL-7B (base) | 35.7 | 0.2263 | 0.8561 | 0.2331 | 16.17 |
| Qwen2.5-VL-7B (SFT) | 27.0 | 0.2297 | 0.8413 | 0.1795 | 19.43 |
| Back2Struct (RL) | 55.0 | 0.4399 | 0.7278 | 0.3830 | 37.83 |
On successfully compiled outputs, the 7B model matches or beats frontier closed models:
| Model | DINO ↑ | LPIPS ↓ | SSIM ↑ | GPT ↑ |
|---|---|---|---|---|
| Gemini-2.5-Pro | 0.8855 | 0.4648 | 0.6532 | 77.95 |
| GPT-5 | 0.8852 | 0.4350 | 0.6573 | 77.21 |
| Back2Struct (7B) | 0.8697 | 0.3896 | 0.6738 | 78.63 |
Usage
Run the model
import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info
model_id = "Pengyu965/Back2Struct-Image2SVG-7B"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
model_id, torch_dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(model_id)
messages = [{"role": "user", "content": [
{"type": "image", "image": "diagram.png"},
{"type": "text", "text": "Convert this structured image into editable SVG/XML code."},
]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(text=[text], images=image_inputs, videos=video_inputs,
padding=True, return_tensors="pt").to(model.device)
generated = model.generate(**inputs, max_new_tokens=8192, do_sample=False)
print(processor.batch_decode(generated[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])
Load the dataset
from datasets import load_dataset
ds = load_dataset("Pengyu965/StructHub")
ex = ds["benchmark"][0]
ex["image"].save("input.png"); print(ex["svg"][:500])
Intended use & limitations
- Intended use: recovering editable vector code from images of structured graphics.
- Not for: photographic / natural-image vectorization.
- Limitations: built on a general (non-code-specialized) VLM; outputs capped at 8,192 tokens; very long/dense diagrams (hard tier) remain the main failure mode (render reliability).
Licensing
Model weights are released under Apache-2.0 (inheriting the Qwen2.5-VL-7B-Instruct base). The StructHub dataset aggregates multiple sources, including crawled open-domain graphics; it is released for research use — review and respect the original content licenses. (Confirm before wide distribution.)
Citation
@misc{yan2026back2structmakingstructuredimages,
title={Back2Struct: Making Structured Images Editable Again},
author={Pengyu Yan and Yixin Wu and Yunjie Tian and David Doermann},
year={2026},
eprint={2609.37016},
archivePrefix={arXiv},
primaryClass={cs.CV},
url={https://arxiv.org/abs/2609.37016},
}
- Downloads last month
- -