GPD-2B

arXiv Code Model Data

Official checkpoint of GPD (Geometry-Privileged Distillation) from the paper Distilling Routed 3D Privilege for Spatial Reasoning in Vision-Language Models.

  • Base model: Qwen3-VL-2B-Instruct
  • Training data: Mixed-15k (VSI 10k + SPAR 4k + MindCube 1k), text_routed variant
  • Training: GRPO + privileged KL on incorrect trajectories (kl_coef=0.003), 111 steps
  • Inference: RGB-only, same interface as the base model

Results

Model VSI-Bench Avg. (MindCube / SPAR / MMSI / ViewSpatial)
Qwen3-VL-2B 51.3 31.1
GRPO 51.3 32.7
OPSD (answer privilege) 50.8 30.2
GPD 51.7 33.6

Usage

from transformers import AutoProcessor, Qwen3VLForConditionalGeneration

model_id = "xinyili0624/GPD-2B"
model = Qwen3VLForConditionalGeneration.from_pretrained(model_id, torch_dtype="auto", device_map="auto")
processor = AutoProcessor.from_pretrained(model_id)

Usage is identical to Qwen3-VL-2B-Instruct; see the base model card for image / video inference examples.

Downloads last month
5
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xinyili0624/GPD-2B

Finetuned
(271)
this model

Paper for xinyili0624/GPD-2B