Instructions to use oofwite/OMG-VLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use oofwite/OMG-VLM with PEFT:
from peft import PeftModel from transformers import AutoModelForCausalLM base_model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen-VL-Chat") model = PeftModel.from_pretrained(base_model, "oofwite/OMG-VLM") - Notebooks
- Google Colab
- Kaggle
OMG-VLM
One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models
Accepted at EMNLP 2026 · Main Conference
Jiayi Yang† · Yifang Chen† · Yuanfu Sun · Jiajin Liu · Qiaoyu Tan
† Equal contribution
One shared vision-language backbone for text-attributed, image-attributed, and multimodal graphs.
This repository holds the fine-tuned OMG-VLM weights: a checkpoint with 32 visual queries and 8 text context tokens, for use with the code at github.com/Jo-eyang/OMG-VLM.
Overview
OMG-VLM is a unified vision-language framework for attributed graph learning under heterogeneous modality schemas. It builds on Qwen-VL and adds graph-aware adapters that inject neighborhood signals into the VLM-native embedding space for both image and text attributes.
| Item | Description |
|---|---|
| Task family | Attributed graph learning with image, text, or mixed node attributes |
| Backbone | Qwen-VL-Chat |
| Graph modules | Graph-aware visual adapter and target-aware textual aggregation |
| This checkpoint | LoRA adapter + both graph modules |
- Graph-aware visual adapter: center image features attend to compressed neighbor visual features in the Qwen-VL embedding space.
- Target-aware textual aggregation: target node text retrieves and compresses neighbor text into learnable
<nbr>context tokens.
Files
| File | Contents |
|---|---|
adapter_model.safetensors, adapter_config.json |
LoRA adapter for the Qwen-VL-Chat language model (r=64, α=16; attn.c_attn, attn.c_proj, w1, w2) |
visual_adapter.pth |
Graph-aware visual adapter (transformer.visual_adapter) |
textual_aggregation.pth |
Target-aware textual aggregation (transformer.textual_aggregation) |
The base Qwen-VL-Chat weights are not included. Download them from Qwen/Qwen-VL-Chat.
Configuration
These values must match at inference time:
| Setting | Value |
|---|---|
| Visual adapter layers / queries / heads | 1 / 32 / 32 |
| Visual compressor | on |
| Textual aggregation heads / pool layers / MLP ratio | 16 / 1 / 4.0 |
| Textual aggregation context tokens | 8 |
| Max neighbors | 10 |
Usage
This checkpoint is loaded with the evaluation script from the OMG-VLM repository.
# 1. Code
git clone https://github.com/Jo-eyang/OMG-VLM.git
cd OMG-VLM
pip install -r requirements.txt
# 2. Base weights go into the repo's Qwen_VL_Chat/ folder, next to the OMG-VLM model code.
# Download only the weight shards so the repo's config and code are kept.
hf download Qwen/Qwen-VL-Chat --local-dir Qwen_VL_Chat \
--include "pytorch_model-*.bin" "pytorch_model.bin.index.json"
# 3. This checkpoint
hf download oofwite/OMG-VLM --local-dir checkpoints/omg_vlm
# 4. Evaluate (the repo root must be on PYTHONPATH so the model code can import omg_vlm)
export PYTHONPATH=$(pwd):$PYTHONPATH
python evaluate_omg_vlm.py \
--model_name_or_path Qwen_VL_Chat \
--adapter_path checkpoints/omg_vlm \
--data_path /path/to/test.json \
--neighbor_data_path /path/to/neighbors.json \
--text_info_path /path/to/text_info.json \
--output_path outputs/predictions.jsonl \
--max_neighbors 10 \
--visual_adapter_num_layers 1 --visual_adapter_num_queries 32 --visual_adapter_num_heads 32 \
--textual_aggregation_num_heads 16 --textual_aggregation_pool_layers 1 \
--textual_aggregation_pool_mlp_ratio 4.0 --textual_aggregation_context_tokens 8
For the conversation, neighbor, and text-attribute file formats, and for training your own checkpoint, see the Data Format and Training sections of the GitHub README.
Scoring note: evaluate_omg_vlm.py reports strict string exact match, so for example yes against a target of yes. counts as wrong. For yes/no or label tasks, normalize case and punctuation before scoring.
Tested with torch 2.7.1, transformers 4.37.2, and peft 0.10.0.
Training
- Base model: Qwen-VL-Chat; the vision encoder is frozen.
- Data: a mixture of node-classification and link-prediction tasks over image, text, and multimodal graphs:
- Amazon Arts (node classification);
- arXiv (node classification);
- RedditS (link prediction, image-only);
- Movies (link prediction).
- Optimization: 3 epochs; learning rate 1e-5 with cosine schedule and 1% warmup; weight decay 0.1; effective batch size 8; max sequence length 2048.
License
These weights are derived from Qwen-VL-Chat and are subject to the Tongyi Qianwen License Agreement.
Citation
@inproceedings{yang2026one,
title={One Model, Many Graphs: Learning over Attributed Graphs across Heterogeneous Modalities with Vision-Language Models},
author={Yang, Jiayi and Chen, Yifang and Sun, Yuanfu and Liu, Jiajin and Tan, Qiaoyu},
booktitle={Proceedings of the 2026 Conference on Empirical Methods in Natural Language Processing},
year={2026}
}
Acknowledgment
This work builds on Qwen-VL and widely used open-source tooling from the VLM/LLM ecosystem. We thank the authors and maintainers of Qwen-VL, Hugging Face Transformers, PEFT, FastChat, DeepSpeed, and related projects.
- Downloads last month
- 7
Model tree for oofwite/OMG-VLM
Base model
Qwen/Qwen-VL-Chat