--- license: cc-by-nc-4.0 library_name: transformers pipeline_tag: image-text-to-text language: - en datasets: - ChartGalaxyPP/ChartGalaxyPlusPlus tags: - chart - infographic - scene-graph - image2scenegraph - qwen3_5 ---
From infographic images to structured chart representations
Dataset · Project · Run inference · Output format
**Image2SceneGraph identifies chart elements and how they belong together.** Fine-tuned from the **Qwen3.5-4B family** on ChartGalaxy++, it predicts visual elements, recognized text, bounding boxes, appearance attributes, semantic groups, and parent links as structured JSON.  *Paper examples: separating nearby visual elements and placing marks in the correct semantic group. [Enlarge](assets/image2scenegraph-results.png)* ## Results Evaluation on **1,000 infographic charts: 500 real and 500 synthetic**. | Model | Node F1 | Hierarchy F1 | Spatial F1 | | --- | ---: | ---: | ---: | | GPT-6 Astra | 76.0% | 50.3% | 72.5% | | **Image2SceneGraph (ours)** | **89.4%** | **85.1%** | **88.0%** | [Full comparison with 10 baselines](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/blob/main/applications/image_to_scene_graph/paper_table.json) · [Metric definitions](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/blob/main/applications/image_to_scene_graph/metric_definitions.json) · [Saved predictions and scores](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/releases/download/v1.0/image-to-scene-graph.tar.gz) These are the paper's scores, including its output-recovery and graph-evaluation procedures. Spatial relationships are derived from the predicted structure and geometry; they are **not directly generated** by this model. ## Download and run ```python from huggingface_hub import HfApi, snapshot_download repo = "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph" revision = HfApi().model_info(repo).sha snapshot_download(repo, revision=revision, local_dir="image2scenegraph-model") ``` Follow **[the complete single-image inference example](INFERENCE.md)** for the prompt, image preprocessing, and generation settings. Keep the weight shards, tokenizer, processor, and chat template together. The tested runtime uses Python 3.11, vLLM 0.20.2, Transformers 5.12.1, and PyTorch 2.11.0 / CUDA 13.0 on one RTX PRO 6000 GPU; exact dependencies are in [requirements-smoke.txt](requirements-smoke.txt). ## Model output | Output | Contents | | --- | --- | | **Elements** | Text, image, and shape items with bounding boxes and appearance descriptions | | **Text** | Recognized strings associated with their visual elements | | **Groups** | Semantic units and parent links organizing the chart | | **Coordinates** | Native `elements.layout` boxes use **`[x0, y0, x1, y1]`** on a 0–1000 grid | **The dataset uses a different serialization:** `compositional_deconstruction.nodes` with **`[y0, x0, y1, x1]`** boxes. Use the mapping in [OUTPUT_FORMAT.md](OUTPUT_FORMAT.md) when comparing outputs with dataset annotations. ## Checkpoint | Property | Released artifact | | --- | --- | | Architecture | `Qwen3_5ForConditionalGeneration` | | Fine-tuning | Supervised fine-tuning for chart scene graph prediction | | Checkpoint | Step 100,000; full model, not an adapter | | Precision | BF16 | | Weights | Two Safetensors shards; tokenizer, processor, and chat template included |