Instructions to use cqqq/GSR-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use cqqq/GSR-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="cqqq/GSR-8B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("cqqq/GSR-8B") model = AutoModelForMultimodalLM.from_pretrained("cqqq/GSR-8B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use cqqq/GSR-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "cqqq/GSR-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cqqq/GSR-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/cqqq/GSR-8B
- SGLang
How to use cqqq/GSR-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "cqqq/GSR-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cqqq/GSR-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "cqqq/GSR-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "cqqq/GSR-8B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use cqqq/GSR-8B with Docker Model Runner:
docker model run hf.co/cqqq/GSR-8B
GSR-8B
GSR-8B is the full-parameter fine-tuned model released with Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection (CVPR 2026). It is trained on paired top-view and side-view X-ray images from GSXray to perform geometric and semantic cross-modal reasoning.
Model details
- Base model:
Qwen/Qwen3-VL-8B-Instruct - Reproducibility revision:
0c351dd01ed87e9c1b53cbc748cba10e6187ff3b - Fine-tuning method: full-parameter supervised fine-tuning
- Training data:
cqqq/GSXray, 44,019 paired-image examples - Framework: LLaMA-Factory at commit
56f45e826f828e44fcdca6a1a5a854d4b71f6ec7 - License: Apache-2.0
The exact revision used by the original local base-model download was not recorded. The revision above is the immutable public reproduction pin selected from the upstream history; it is not presented as a recovered download record.
The model adds paired special tokens for top, side, think, conclusion,
answer, bbox, and label. For reproducible preprocessing and prompting,
load the processor and chat template distributed in this repository.
Quick start
from transformers import AutoProcessor, Qwen3VLForConditionalGeneration
repo_id = "cqqq/GSR-8B"
model = Qwen3VLForConditionalGeneration.from_pretrained(
repo_id,
torch_dtype="auto",
device_map="auto",
)
processor = AutoProcessor.from_pretrained(repo_id)
Pass the two X-ray views in the same top-view then side-view order used by the GSXray records. The public GSR code repository provides the complete prompt, training, serving, and evaluation entry points.
Training
The preserved, path-free hyperparameters are in training_config.yaml; the
recorded final Trainer metrics are in training_summary.json. The release
contains inference weights only. Optimizer state, scheduler state, RNG state,
intermediate checkpoints, TensorBoard logs, and machine-local paths are
intentionally excluded.
Intended use and limitations
This research model is intended for reproducible evaluation of dual-view X-ray reasoning. It can produce incorrect descriptions, locations, or conclusions and must not be treated as an autonomous security decision system. Results can depend on image acquisition conditions, prompts, decoding settings, and data distribution. Human review and application-specific validation are required.
Integrity
checksums.sha256 covers every published file except the checksum file itself.
release_manifest.json records the closed release file set and provenance pins.
Citation
@inproceedings{peng2026gsr,
title = {Can a Second-View Image Be a Language? Geometric and Semantic Cross-Modal Reasoning for X-ray Prohibited Item Detection},
author = {Peng, Chuang and Tao, Renshuai and Ren, Zhongwei and Liu, Xianglong and Wei, Yunchao},
booktitle = {Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition},
year = {2026}
}
Attribution
GSR-8B is a modified derivative of Qwen3-VL-8B-Instruct. See NOTICE for
copyright and upstream attribution.
- Downloads last month
- 27
Model tree for cqqq/GSR-8B
Base model
Qwen/Qwen3-VL-8B-Instruct