SPARC-Qwen3.5-4B

Qwen3.5-4B fully fine-tuned for embodied spatial reasoning using VQA data generated from SPARC annotations.

Training data

The training mixture contains SPARC-generated VQA data from ours_adaptive_det_soft_snr_sp8, FSD, RoboPoint, and LLaVA-OneVision2. SPARC samples use an annotation-quality threshold of 0.97, are sorted by score, and are capped at 700 samples per object. The paper reports 1,159,047 training pairs for this mixture.

Release SPARC VQA (filtered) FSD RoboPoint LLaVA-OneVision2 EO-1.5M
Qwen3.5-4B Yes Yes Yes Yes No
Qwen3.5-0.8B-VTFT Yes Yes Yes Yes No
Qwen3.5-9B-EO Yes Yes Yes Yes Yes

The SPARC training data is the filtered release subset, which needs no further SPARC filtering. Its filtering script and release mixture manifest are in irl-kit/SPARC-VQA-Raw. FSD, RoboPoint, LLaVA-OneVision2, and EO-1.5M remain their respective upstream datasets.

Prompting

Prompt formatting is important for these models. Use the bundled chat_template.jinja through processor.apply_chat_template(..., add_generation_prompt=True), with one user turn containing the image(s) followed by the text question. Disable thinking/reasoning mode to match evaluation.

For a single point, append exactly:

Output the point coordinates in JSON format like [{"point_2d": [x, y], "label": "target"}]. Use integer coordinates between 0 and 1000.

For a trajectory or multiple points, append exactly:

Return only a JSON list like [{"point_2d": [x1, y1], "label": "point_1"}, {"point_2d": [x2, y2], "label": "point_2"}, ...]. Use integer coordinates between 0 and 1000.

Training

The vision encoder is frozen and the vision projector is trainable. The model was fully fine-tuned for one epoch with a learning rate of 2e-5 and a maximum sequence length of 5600.

Evaluation

This is the paper's primary 4B model. The paper reports a 62.7 pointing/VQA average for this model. The local full benchmark evaluation associates the released weights with aggregate score 0.698.

Model Aggregate Where2Place RefSpatial location IA-Bench RoboRefIt testA VA Bench-P
Qwen3.5-4B 0.698 72.0 59.0 79.0 85.7 65.7
Qwen3.5-0.8B-VTFT 0.605 58.0 47.0 76.7 80.9 48.3
Qwen3.5-9B-EO 0.719 76.0 68.0 78.5 85.2 68.7

Citation

@article{blank2026sparc,
  title={SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale},
  author={Blank, Nils and others},
  journal={arXiv preprint arXiv:2606.13497},
  year={2026}
}
Downloads last month
26
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train irl-kit/SPARC-Qwen3.5-4B

Collection including irl-kit/SPARC-Qwen3.5-4B

Paper for irl-kit/SPARC-Qwen3.5-4B