Back2Struct: Making Structured Images Editable Again

Recovering editable, object-level SVG / XML code from images of structured graphics.

Back2Struct "makes structured images editable again" by directly recovering vector graphics code (SVG / XML) from image representations. Given an image of a structured graphic — a diagram, chart, flowchart, architecture figure, UML diagram, or schema — it predicts semantically object-level SVG/XML that explicitly encodes text, shapes, topology, and layout (not low-level pixel tracing), so the code imports straight into tools like PowerPoint to edit, restyle, and reuse while preserving structure. Training is supervised fine-tuning from Qwen2.5-VL-7B-Instruct, followed by GRPO with a composite reward (syntactic validity, length fidelity, perceptual quality).

This project ships two artifacts — a model and a dataset — documented together in this shared card.

🤖 Model — Back2Struct-Image2SVG-7B

Base Qwen/Qwen2.5-VL-7B-Instruct (7B)
Stage 1 — SFT Supervised fine-tuning on StructHub image→SVG pairs
Stage 2 — GRPO Group-Relative Policy Optimization (TRL)
Reward Composite, render-gated: syntactic validity (must render) × perceptual fidelity (baseline-subtracted DINOv2 similarity between the rendered prediction and the target), encouraging compilability, length consistency, and structural/semantic faithfulness
Max output 8,192 tokens · LoRA (r=64, α=128), merged into the released weights

📚 Dataset — StructHub

Structured images paired with their editable SVG/XML source code.

Split Description # examples
train Full training set (all sources, all token lengths), benchmark held out 86,102
benchmark Held-out evaluation benchmark (starvector 569 + crawled 334 + nn_diagram 97) 1,000

Fields: id (string), source (string), token_count (int32), image (embedded PNG), svg (ground-truth SVG/XML). The benchmark is stratified by SVG length into easy (≤2,048), medium (2,048–4,096), and hard (>4,096) tiers (333/333/334).

Results

StructHub benchmark, render-gated (a non-rendering prediction scores 0 on DINO/SSIM/GPT, 1 on LPIPS). SR = render success (%); GPT = 0–100 GPTScore.

Model SR ↑ DINO ↑ LPIPS ↓ SSIM ↑ GPT ↑
Overall
Qwen3-VL-30B-A3B 18.8 0.1566 0.9067 0.1201 13.29
StarVector-8B 31.4 0.2694 0.7936 0.2170 21.05
Qwen2.5-VL-7B (base) 52.3 0.3527 0.7834 0.3276 27.56
Qwen2.5-VL-7B (SFT) 42.8 0.3715 0.7403 0.2891 33.49
Back2Struct (RL) 71.5 0.5622 0.6473 0.4966 51.85
Easy
Qwen3-VL-30B-A3B 26.5 0.2231 0.8659 0.1702 19.51
StarVector-8B 46.1 0.3923 0.7023 0.3123 31.90
Qwen2.5-VL-7B (base) 69.3 0.4776 0.7123 0.4196 39.66
Qwen2.5-VL-7B (SFT) 56.0 0.4866 0.6630 0.3789 44.58
Back2Struct (RL) 84.0 0.6384 0.5956 0.5796 62.63
Medium
Qwen3-VL-30B-A3B 19.6 0.1631 0.9004 0.1288 13.38
StarVector-8B 37.0 0.3232 0.7533 0.2583 25.27
Qwen2.5-VL-7B (base) 51.8 0.3546 0.7816 0.3306 26.88
Qwen2.5-VL-7B (SFT) 45.5 0.3985 0.7161 0.3091 36.50
Back2Struct (RL) 75.6 0.6088 0.6182 0.5275 55.14
Hard
Qwen3-VL-30B-A3B 10.2 0.0836 0.9536 0.0614 7.01
StarVector-8B 11.1 0.0932 0.9249 0.0808 6.02
Qwen2.5-VL-7B (base) 35.7 0.2263 0.8561 0.2331 16.17
Qwen2.5-VL-7B (SFT) 27.0 0.2297 0.8413 0.1795 19.43
Back2Struct (RL) 55.0 0.4399 0.7278 0.3830 37.83

On successfully compiled outputs, the 7B model matches or beats frontier closed models:

Model DINO ↑ LPIPS ↓ SSIM ↑ GPT ↑
Gemini-2.5-Pro 0.8855 0.4648 0.6532 77.95
GPT-5 0.8852 0.4350 0.6573 77.21
Back2Struct (7B) 0.8697 0.3896 0.6738 78.63

Usage

Run the model

import torch
from transformers import Qwen2_5_VLForConditionalGeneration, AutoProcessor
from qwen_vl_utils import process_vision_info

model_id = "Pengyu965/Back2Struct-Image2SVG-7B"
model = Qwen2_5_VLForConditionalGeneration.from_pretrained(
    model_id, torch_dtype=torch.bfloat16, device_map="auto")
processor = AutoProcessor.from_pretrained(model_id)

messages = [{"role": "user", "content": [
    {"type": "image", "image": "diagram.png"},
    {"type": "text",  "text": "Convert this structured image into editable SVG/XML code."},
]}]
text = processor.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
image_inputs, video_inputs = process_vision_info(messages)
inputs = processor(text=[text], images=image_inputs, videos=video_inputs,
                   padding=True, return_tensors="pt").to(model.device)
generated = model.generate(**inputs, max_new_tokens=8192, do_sample=False)
print(processor.batch_decode(generated[:, inputs.input_ids.shape[1]:], skip_special_tokens=True)[0])

Load the dataset

from datasets import load_dataset
ds = load_dataset("Pengyu965/StructHub")
ex = ds["benchmark"][0]
ex["image"].save("input.png"); print(ex["svg"][:500])

Intended use & limitations

  • Intended use: recovering editable vector code from images of structured graphics.
  • Not for: photographic / natural-image vectorization.
  • Limitations: built on a general (non-code-specialized) VLM; outputs capped at 8,192 tokens; very long/dense diagrams (hard tier) remain the main failure mode (render reliability).

Licensing

Model weights are released under Apache-2.0 (inheriting the Qwen2.5-VL-7B-Instruct base). The StructHub dataset aggregates multiple sources, including crawled open-domain graphics; it is released for research use — review and respect the original content licenses. (Confirm before wide distribution.)

Citation

@misc{yan2026back2structmakingstructuredimages,
      title={Back2Struct: Making Structured Images Editable Again},
      author={Pengyu Yan and Yixin Wu and Yunjie Tian and David Doermann},
      year={2026},
      eprint={2609.37016},
      archivePrefix={arXiv},
      primaryClass={cs.CV},
      url={https://arxiv.org/abs/2609.37016},
}
Downloads last month
-
Safetensors
Model size
8B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Pengyu965/Back2Struct-Image2SVG-7B

Finetuned
(1239)
this model
Quantizations
1 model

Dataset used to train Pengyu965/Back2Struct-Image2SVG-7B

Paper for Pengyu965/Back2Struct-Image2SVG-7B