rsoohyun's picture
Create README.md
d8af634 verified
|
Raw
History Blame Contribute Delete
976 Bytes
metadata
license: apache-2.0
library_name: transformers
pipeline_tag: image-text-to-text
datasets:
  - rsoohyun/SpatialBlock-15k
base_model:
  - Qwen/Qwen2.5-VL-3B-Instruct

SpatialBlock-3B-direct

This repository contains the SpatialBlock-3B-direct checkpoint from the paper SpatialBlock: Enhancing Spatial Intelligence in LVLMs via Synthetic Block-Stacking Problem.

It is a fine-tuned version of Qwen2.5-VL-3B-Instruct on the synthetic SpatialBlock-15k dataset. The model directly predicts answers to spatial reasoning tasks such as 3D-to-2D projection, viewpoint transformation, and structural combination.

For training details, evaluation results, and the companion “reason” model, please refer to the GitHub repository: https://github.com/rsoohyun/SpatialBlock.