SPARC-Qwen3.5-9B-EO

Qwen3.5-9B fully fine-tuned for embodied spatial reasoning using SPARC-generated VQA data and an EO-1.5M-augmented mixture.

Training data

The training mixture contains SPARC-generated VQA data from ours_adaptive_det_soft_snr_sp8, FSD, RoboPoint, LLaVA-OneVision2, and EO-1.5M. SPARC samples use an annotation-quality threshold of 0.97, are sorted by score, and are capped at 700 samples per object.

Release SPARC VQA (filtered) FSD RoboPoint LLaVA-OneVision2 EO-1.5M
Qwen3.5-4B Yes Yes Yes Yes No
Qwen3.5-0.8B-VTFT Yes Yes Yes Yes No
Qwen3.5-9B-EO Yes Yes Yes Yes Yes

The SPARC training data is the filtered release subset, which needs no further SPARC filtering. Its filtering script and release mixture manifest are in irl-kit/SPARC-VQA-Raw. FSD, RoboPoint, LLaVA-OneVision2, and EO-1.5M remain their respective upstream datasets.

Prompting

Prompt formatting is important for these models. Use the bundled chat_template.jinja through processor.apply_chat_template(..., add_generation_prompt=True), with one user turn containing the image(s) followed by the text question. Disable thinking/reasoning mode to match evaluation.

For a single point, append exactly:

Output the point coordinates in JSON format like [{"point_2d": [x, y], "label": "target"}]. Use integer coordinates between 0 and 1000.

For a trajectory or multiple points, append exactly:

Return only a JSON list like [{"point_2d": [x1, y1], "label": "point_1"}, {"point_2d": [x2, y2], "label": "point_2"}, ...]. Use integer coordinates between 0 and 1000.

Training

The vision encoder is frozen and the vision projector is trainable. The model was fully fine-tuned for one epoch with a learning rate of 2e-5 and a maximum sequence length of 5600.

Evaluation

This is the strongest evaluated 9B run: local full-benchmark aggregate 0.719. It should be compared separately from SPARC-only mixtures because it includes EO-1.5M.

Model Aggregate Where2Place RefSpatial location IA-Bench RoboRefIt testA VA Bench-P
Qwen3.5-4B 0.698 72.0 59.0 79.0 85.7 65.7
Qwen3.5-0.8B-VTFT 0.605 58.0 47.0 76.7 80.9 48.3
Qwen3.5-9B-EO 0.719 76.0 68.0 78.5 85.2 68.7

Citation

@article{blank2026sparc,
  title={SPARC: Reliable Spatial Annotations from Robot Demonstrations at Scale},
  author={Blank, Nils and others},
  journal={arXiv preprint arXiv:2606.13497},
  year={2026}
}
Downloads last month
13
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for irl-kit/SPARC-Qwen3.5-9B-EO