--- license: cc-by-nc-4.0 library_name: transformers pipeline_tag: image-text-to-text language: - en datasets: - ChartGalaxyPP/ChartGalaxyPlusPlus tags: - chart - infographic - scene-graph - image2scenegraph - qwen3_5 ---

ChartGalaxy++ · Image2SceneGraph

From infographic images to structured chart representations

Dataset   ·   Project   ·   Run inference   ·   Output format

**Image2SceneGraph identifies chart elements and how they belong together.** Fine-tuned from the **Qwen3.5-4B family** on ChartGalaxy++, it predicts visual elements, recognized text, bounding boxes, appearance attributes, semantic groups, and parent links as structured JSON. ![Qualitative comparisons from the paper showing element localization and grouping by Image2SceneGraph and GPT-6 Astra](assets/image2scenegraph-results.png) *Paper examples: separating nearby visual elements and placing marks in the correct semantic group. [Enlarge](assets/image2scenegraph-results.png)* ## Results Evaluation on **1,000 infographic charts: 500 real and 500 synthetic**. | Model | Node F1 | Hierarchy F1 | Spatial F1 | | --- | ---: | ---: | ---: | | GPT-6 Astra | 76.0% | 50.3% | 72.5% | | **Image2SceneGraph (ours)** | **89.4%** | **85.1%** | **88.0%** | [Full comparison with 10 baselines](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/blob/main/applications/image_to_scene_graph/paper_table.json) · [Metric definitions](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/blob/main/applications/image_to_scene_graph/metric_definitions.json) · [Saved predictions and scores](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/releases/download/v1.0/image-to-scene-graph.tar.gz) These are the paper's scores, including its output-recovery and graph-evaluation procedures. Spatial relationships are derived from the predicted structure and geometry; they are **not directly generated** by this model. ## Download and run ```python from huggingface_hub import HfApi, snapshot_download repo = "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph" revision = HfApi().model_info(repo).sha snapshot_download(repo, revision=revision, local_dir="image2scenegraph-model") ``` Follow **[the complete single-image inference example](INFERENCE.md)** for the prompt, image preprocessing, and generation settings. Keep the weight shards, tokenizer, processor, and chat template together. The tested runtime uses Python 3.11, vLLM 0.20.2, Transformers 5.12.1, and PyTorch 2.11.0 / CUDA 13.0 on one RTX PRO 6000 GPU; exact dependencies are in [requirements-smoke.txt](requirements-smoke.txt). ## Model output | Output | Contents | | --- | --- | | **Elements** | Text, image, and shape items with bounding boxes and appearance descriptions | | **Text** | Recognized strings associated with their visual elements | | **Groups** | Semantic units and parent links organizing the chart | | **Coordinates** | Native `elements.layout` boxes use **`[x0, y0, x1, y1]`** on a 0–1000 grid | **The dataset uses a different serialization:** `compositional_deconstruction.nodes` with **`[y0, x0, y1, x1]`** boxes. Use the mapping in [OUTPUT_FORMAT.md](OUTPUT_FORMAT.md) when comparing outputs with dataset annotations. ## Checkpoint | Property | Released artifact | | --- | --- | | Architecture | `Qwen3_5ForConditionalGeneration` | | Fine-tuning | Supervised fine-tuning for chart scene graph prediction | | Checkpoint | Step 100,000; full model, not an adapter | | Precision | BF16 | | Weights | Two Safetensors shards; tokenizer, processor, and chat template included |
Validation and limitations The export matches the checkpoint used for the paper's evaluation. File identities are recorded in [release_metadata.json](release_metadata.json) and [SHA256SUMS](SHA256SUMS). A standalone GPU smoke test produced 128 layout items and 11,485 output tokens, stopping normally with valid JSON without recovery. The native schema and graph-reference checks also passed. See [the validation receipt](smoke/validation.json) and [saved prediction](smoke/prediction.json). This checks loading and output structure, not prediction accuracy; its input hash identifies the historical image version. Predictions may omit elements, misread text, assign incorrect parents, or produce invalid JSON. Dense charts may exceed the 20,000-token output limit. Valid JSON alone does not establish completeness or correctness. The family name “4B” is not the count of all scalars in this export, which also contains vision and auxiliary components.
## License ChartGalaxy++ fine-tuning contributions use **CC BY-NC 4.0**. Retained Qwen materials remain subject to their upstream **Apache-2.0** license. See [LICENSE](LICENSE), [NOTICE.md](NOTICE.md), and [the upstream license](LICENSE-QWEN-APACHE-2.0). This repository contains model artifacts and usage documentation. Annotation, training, evaluation, and review pipeline implementations are not included.