OMG-VLM

One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models

Accepted at EMNLP 2026 · Main Conference

Jiayi Yang† · Yifang Chen† · Yuanfu Sun · Jiajin Liu · Qiaoyu Tan
† Equal contribution

Paper Code Qwen-VL

Overview of the OMG-VLM framework

One shared vision-language backbone for text-attributed, image-attributed, and multimodal graphs.

This repository holds the fine-tuned OMG-VLM weights: a checkpoint with 32 visual queries and 8 text context tokens, for use with the code at github.com/Jo-eyang/OMG-VLM.

Overview

OMG-VLM is a unified vision-language framework for attributed graph learning under heterogeneous modality schemas. It builds on Qwen-VL and adds graph-aware adapters that inject neighborhood signals into the VLM-native embedding space for both image and text attributes.

Item Description
Task family Attributed graph learning with image, text, or mixed node attributes
Backbone Qwen-VL-Chat
Graph modules Graph-aware visual adapter and target-aware textual aggregation
This checkpoint LoRA adapter + both graph modules
  • Graph-aware visual adapter: center image features attend to compressed neighbor visual features in the Qwen-VL embedding space.
  • Target-aware textual aggregation: target node text retrieves and compresses neighbor text into learnable <nbr> context tokens.

Files

File Contents
adapter_model.safetensors, adapter_config.json LoRA adapter for the Qwen-VL-Chat language model (r=64, α=16; attn.c_attn, attn.c_proj, w1, w2)
visual_adapter.pth Graph-aware visual adapter (transformer.visual_adapter)
textual_aggregation.pth Target-aware textual aggregation (transformer.textual_aggregation)

The base Qwen-VL-Chat weights are not included. Download them from Qwen/Qwen-VL-Chat.

Configuration

These values must match at inference time:

Setting Value
Visual adapter layers / queries / heads 1 / 32 / 32
Visual compressor on
Textual aggregation heads / pool layers / MLP ratio 16 / 1 / 4.0
Textual aggregation context tokens 8
Max neighbors 10

Usage

This checkpoint is loaded with the evaluation script from the OMG-VLM repository.

# 1. Code
git clone https://github.com/Jo-eyang/OMG-VLM.git
cd OMG-VLM
pip install -r requirements.txt

# 2. Base weights go into the repo's Qwen_VL_Chat/ folder, next to the OMG-VLM model code.
#    Download only the weight shards so the repo's config and code are kept.
hf download Qwen/Qwen-VL-Chat --local-dir Qwen_VL_Chat \
    --include "pytorch_model-*.bin" "pytorch_model.bin.index.json"

# 3. This checkpoint
hf download oofwite/OMG-VLM --local-dir checkpoints/omg_vlm

# 4. Evaluate (the repo root must be on PYTHONPATH so the model code can import omg_vlm)
export PYTHONPATH=$(pwd):$PYTHONPATH
python evaluate_omg_vlm.py \
  --model_name_or_path Qwen_VL_Chat \
  --adapter_path checkpoints/omg_vlm \
  --data_path /path/to/test.json \
  --neighbor_data_path /path/to/neighbors.json \
  --text_info_path /path/to/text_info.json \
  --output_path outputs/predictions.jsonl \
  --max_neighbors 10 \
  --visual_adapter_num_layers 1 --visual_adapter_num_queries 32 --visual_adapter_num_heads 32 \
  --textual_aggregation_num_heads 16 --textual_aggregation_pool_layers 1 \
  --textual_aggregation_pool_mlp_ratio 4.0 --textual_aggregation_context_tokens 8

For the conversation, neighbor, and text-attribute file formats, and for training your own checkpoint, see the Data Format and Training sections of the GitHub README.

Scoring note: evaluate_omg_vlm.py reports strict string exact match, so for example yes against a target of yes. counts as wrong. For yes/no or label tasks, normalize case and punctuation before scoring.

Tested with torch 2.7.1, transformers 4.37.2, and peft 0.10.0.

Training

  • Base model: Qwen-VL-Chat; the vision encoder is frozen.
  • Data: a mixture of node-classification and link-prediction tasks over image, text, and multimodal graphs:
    • Amazon Arts (node classification);
    • arXiv (node classification);
    • RedditS (link prediction, image-only);
    • Movies (link prediction).
  • Optimization: 3 epochs; learning rate 1e-5 with cosine schedule and 1% warmup; weight decay 0.1; effective batch size 8; max sequence length 2048.

License

These weights are derived from Qwen-VL-Chat and are subject to the Tongyi Qianwen License Agreement.

Citation

@inproceedings{yang2026one,
  title={One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models},
  author={Yang, Jiayi and Chen, Yifang and Sun, Yuanfu and Liu, Jiajin and Tan, Qiaoyu},
  booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
  year={2026}
}

Acknowledgment

This work builds on Qwen-VL and widely used open-source tooling from the VLM/LLM ecosystem. We thank the authors and maintainers of Qwen-VL, Hugging Face Transformers, PEFT, FastChat, DeepSpeed, and related projects.

Downloads last month
7
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for oofwite/OMG-VLM

Adapter
(47)
this model

Paper for oofwite/OMG-VLM