How to use from
Docker Model Runner
docker model run hf.co/ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph
Quick Links

ChartGalaxy++ · Image2SceneGraph

From infographic images to structured chart representations

Dataset   ·   Project   ·   Run inference   ·   Output format

Image2SceneGraph identifies chart elements and how they belong together. Fine-tuned from the Qwen3.5-4B family on ChartGalaxy++, it predicts visual elements, recognized text, bounding boxes, appearance attributes, semantic groups, and parent links as structured JSON.

Qualitative comparisons from the paper showing element localization and grouping by Image2SceneGraph and GPT-6 Astra

Paper examples: separating nearby visual elements and placing marks in the correct semantic group. Enlarge

Results

Evaluation on 1,000 infographic charts: 500 real and 500 synthetic.

Model Node F1 Hierarchy F1 Spatial F1
GPT-6 Astra 76.0% 50.3% 72.5%
Image2SceneGraph (ours) 89.4% 85.1% 88.0%

Full comparison with 10 baselines · Metric definitions · Saved predictions and scores

These are the paper's scores, including its output-recovery and graph-evaluation procedures. Spatial relationships are derived from the predicted structure and geometry; they are not directly generated by this model.

Download and run

from huggingface_hub import HfApi, snapshot_download

repo = "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph"
revision = HfApi().model_info(repo).sha
snapshot_download(repo, revision=revision, local_dir="image2scenegraph-model")

Follow the complete single-image inference example for the prompt, image preprocessing, and generation settings. Keep the weight shards, tokenizer, processor, and chat template together. The tested runtime uses Python 3.11, vLLM 0.20.2, Transformers 5.12.1, and PyTorch 2.11.0 / CUDA 13.0 on one RTX PRO 6000 GPU; exact dependencies are in requirements-smoke.txt.

Model output

Output Contents
Elements Text, image, and shape items with bounding boxes and appearance descriptions
Text Recognized strings associated with their visual elements
Groups Semantic units and parent links organizing the chart
Coordinates Native elements.layout boxes use [x0, y0, x1, y1] on a 0–1000 grid

The dataset uses a different serialization: compositional_deconstruction.nodes with [y0, x0, y1, x1] boxes. Use the mapping in OUTPUT_FORMAT.md when comparing outputs with dataset annotations.

Checkpoint

Property Released artifact
Architecture Qwen3_5ForConditionalGeneration
Fine-tuning Supervised fine-tuning for chart scene graph prediction
Checkpoint Step 100,000; full model, not an adapter
Precision BF16
Weights Two Safetensors shards; tokenizer, processor, and chat template included
Validation and limitations

The export matches the checkpoint used for the paper's evaluation. File identities are recorded in release_metadata.json and SHA256SUMS.

A standalone GPU smoke test produced 128 layout items and 11,485 output tokens, stopping normally with valid JSON without recovery. The native schema and graph-reference checks also passed. See the validation receipt and saved prediction. This checks loading and output structure, not prediction accuracy; its input hash identifies the historical image version.

Predictions may omit elements, misread text, assign incorrect parents, or produce invalid JSON. Dense charts may exceed the 20,000-token output limit. Valid JSON alone does not establish completeness or correctness. The family name “4B” is not the count of all scalars in this export, which also contains vision and auxiliary components.

License

ChartGalaxy++ fine-tuning contributions use CC BY-NC 4.0. Retained Qwen materials remain subject to their upstream Apache-2.0 license. See LICENSE, NOTICE.md, and the upstream license.

This repository contains model artifacts and usage documentation. Annotation, training, evaluation, and review pipeline implementations are not included.

Downloads last month
14
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph