Image-Text-to-Text
Transformers
Safetensors
English
qwen3_5
chart
infographic
scene-graph
image2scenegraph
conversational
Instructions to use ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph") model = AutoModelForMultimodalLM.from_pretrained("ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph
- SGLang
How to use ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph with Docker Model Runner:
docker model run hf.co/ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph
File size: 5,453 Bytes
1f05665 e3af8b3 1f05665 e3af8b3 1f05665 815d6c5 1f05665 e3af8b3 1f05665 e3af8b3 1f05665 e3af8b3 1f05665 e3af8b3 815d6c5 e3af8b3 815d6c5 e3af8b3 815d6c5 e3af8b3 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 1f05665 815d6c5 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 | ---
license: cc-by-nc-4.0
library_name: transformers
pipeline_tag: image-text-to-text
language:
- en
datasets:
- ChartGalaxyPP/ChartGalaxyPlusPlus
tags:
- chart
- infographic
- scene-graph
- image2scenegraph
- qwen3_5
---
<h1 align="center">ChartGalaxy++ · Image2SceneGraph</h1>
<p align="center"><strong>From infographic images to structured chart representations</strong></p>
<p align="center"><a href="https://huggingface.co/datasets/ChartGalaxyPP/ChartGalaxyPlusPlus">Dataset</a> · <a href="https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus">Project</a> · <a href="https://huggingface.co/ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph/blob/main/INFERENCE.md">Run inference</a> · <a href="https://huggingface.co/ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph/blob/main/OUTPUT_FORMAT.md">Output format</a></p>
**Image2SceneGraph identifies chart elements and how they belong together.** Fine-tuned from the **Qwen3.5-4B family** on ChartGalaxy++, it predicts visual elements, recognized text, bounding boxes, appearance attributes, semantic groups, and parent links as structured JSON.

*Paper examples: separating nearby visual elements and placing marks in the correct semantic group. [Enlarge](assets/image2scenegraph-results.png)*
## Results
Evaluation on **1,000 infographic charts: 500 real and 500 synthetic**.
| Model | Node F1 | Hierarchy F1 | Spatial F1 |
| --- | ---: | ---: | ---: |
| GPT-6 Astra | 76.0% | 50.3% | 72.5% |
| **Image2SceneGraph (ours)** | **89.4%** | **85.1%** | **88.0%** |
[Full comparison with 10 baselines](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/blob/main/applications/image_to_scene_graph/paper_table.json) · [Metric definitions](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/blob/main/applications/image_to_scene_graph/metric_definitions.json) · [Saved predictions and scores](https://github.com/ChartGalaxyPP/ChartGalaxyPlusPlus/releases/download/v1.0/image-to-scene-graph.tar.gz)
These are the paper's scores, including its output-recovery and graph-evaluation procedures. Spatial relationships are derived from the predicted structure and geometry; they are **not directly generated** by this model.
## Download and run
```python
from huggingface_hub import HfApi, snapshot_download
repo = "ChartGalaxyPP/ChartGalaxyPlusPlus-Image2SceneGraph"
revision = HfApi().model_info(repo).sha
snapshot_download(repo, revision=revision, local_dir="image2scenegraph-model")
```
Follow **[the complete single-image inference example](INFERENCE.md)** for the prompt, image preprocessing, and generation settings. Keep the weight shards, tokenizer, processor, and chat template together. The tested runtime uses Python 3.11, vLLM 0.20.2, Transformers 5.12.1, and PyTorch 2.11.0 / CUDA 13.0 on one RTX PRO 6000 GPU; exact dependencies are in [requirements-smoke.txt](requirements-smoke.txt).
## Model output
| Output | Contents |
| --- | --- |
| **Elements** | Text, image, and shape items with bounding boxes and appearance descriptions |
| **Text** | Recognized strings associated with their visual elements |
| **Groups** | Semantic units and parent links organizing the chart |
| **Coordinates** | Native `elements.layout` boxes use **`[x0, y0, x1, y1]`** on a 0–1000 grid |
**The dataset uses a different serialization:** `compositional_deconstruction.nodes` with **`[y0, x0, y1, x1]`** boxes. Use the mapping in [OUTPUT_FORMAT.md](OUTPUT_FORMAT.md) when comparing outputs with dataset annotations.
## Checkpoint
| Property | Released artifact |
| --- | --- |
| Architecture | `Qwen3_5ForConditionalGeneration` |
| Fine-tuning | Supervised fine-tuning for chart scene graph prediction |
| Checkpoint | Step 100,000; full model, not an adapter |
| Precision | BF16 |
| Weights | Two Safetensors shards; tokenizer, processor, and chat template included |
<details>
<summary>Validation and limitations</summary>
The export matches the checkpoint used for the paper's evaluation. File identities are recorded in [release_metadata.json](release_metadata.json) and [SHA256SUMS](SHA256SUMS).
A standalone GPU smoke test produced 128 layout items and 11,485 output tokens, stopping normally with valid JSON without recovery. The native schema and graph-reference checks also passed. See [the validation receipt](smoke/validation.json) and [saved prediction](smoke/prediction.json). This checks loading and output structure, not prediction accuracy; its input hash identifies the historical image version.
Predictions may omit elements, misread text, assign incorrect parents, or produce invalid JSON. Dense charts may exceed the 20,000-token output limit. Valid JSON alone does not establish completeness or correctness. The family name “4B” is not the count of all scalars in this export, which also contains vision and auxiliary components.
</details>
## License
ChartGalaxy++ fine-tuning contributions use **CC BY-NC 4.0**. Retained Qwen materials remain subject to their upstream **Apache-2.0** license. See [LICENSE](LICENSE), [NOTICE.md](NOTICE.md), and [the upstream license](LICENSE-QWEN-APACHE-2.0).
This repository contains model artifacts and usage documentation. Annotation, training, evaluation, and review pipeline implementations are not included.
|