siglip2-base-patch16-512 — ONNX (browser-ready)

ONNX export of google/siglip2-base-patch16-512 (revision a89f5c5), split into a vision and a text model, for onnxruntime-web on WebGPU and wasm (and any other onnxruntime backend). fp32 and fp16, opset 18.

File Precision Size Inputs → output
vision_model.onnx + vision_model.onnx_data fp32 374 MB pixel_values float32 (B, 3, 512, 512) → image_embeds float32 (B, 768)
text_model.onnx + text_model.onnx_data fp32 1.13 GB input_ids int64 (B, 64) → text_embeds float32 (B, 768)
vision_model_fp16.onnx + vision_model_fp16.onnx_data fp16 204 MB same as fp32
text_model_fp16.onnx + text_model_fp16.onnx_data fp16 565 MB same as fp32

B (batch) is dynamic; images are 512×512. Embeddings are not normalised. config.json, preprocessor_config.json and the tokenizer files are copied from the original repo; export_meta.json holds logit_scale (already exponentiated) and logit_bias for scoring.

fp16 variants (*_fp16.onnx) are for onnxruntime-web WebGPU (needs the shader-f16 feature): half the download, faster inference. Inputs and outputs stay float32, so they are drop-in replacements. The vision model keeps its embeddings and attention-pooling head in fp32. WebGPU computes in half precision, so expect small score shifts: in our tests image embeddings matched fp32 with cosine ≥ 0.9995 and label probabilities moved by up to ~2 points. On wasm/CPU use the fp32 files.

Export notes

This is a fixed-resolution SigLIP2 checkpoint (SiglipModel architecture). The vision tower is exported as-is for 512×512 inputs with a dynamic batch dimension; its learned position embeddings are tied to that resolution.

Parity

fp32

Compared with the PyTorch model on onnxruntime (CPUExecutionProvider), threshold cosine ≥ 0.9999: passed.

Tower Inputs Min cosine Max abs diff
vision 4 1.0000000 1.5e-05
text 6 1.0000000 1.8e-05

Max logit difference: 6.0e-05.

Versions: torch 2.14.1, transformers 5.19.0, onnxruntime 1.30.0.

fp16

Compared with the PyTorch model on onnxruntime (CPUExecutionProvider), threshold cosine ≥ 0.999: passed.

Tower Inputs Min cosine Max abs diff
vision 4 0.9999950 9.4e-03
text 6 0.9999990 1.1e-02

Max logit difference: 2.9e-02.

Versions: torch 2.14.1, transformers 5.19.0, onnxruntime 1.30.0.

Inputs

Image. Resize to 512×512 (bilinear, antialiased when downscaling; the aspect ratio is not kept), rescale to [-1, 1] (x / 127.5 - 1) and lay out channels-first as (3, 512, 512). Use SiglipImageProcessor from transformers as the reference.

Text. Lowercase the text, tokenise with the included Gemma tokenizer.json, append <eos> and pad with <pad> (id 0) to exactly 64 tokens. The model pools the last token, so the padding is required. In Python, use Siglip2Tokenizer (transformers 5.x AutoTokenizer returns a GemmaTokenizer that does not lowercase).

Scoring. p = sigmoid(logit_scale · cos(image, text) + logit_bias), per label. SigLIP scores labels independently, so probabilities do not sum to 1; a correct label at 10–40 % with the rest near 0 is normal.

Usage with onnxruntime-web

import * as ort from 'onnxruntime-web/webgpu';

const base = 'https://huggingface.co/upshift/siglip2-base-patch16-512-onnx/resolve/main/';
const vision = await ort.InferenceSession.create(base + 'vision_model.onnx', {
  executionProviders: ['webgpu', 'wasm'],
  externalData: [{ path: 'vision_model.onnx_data', data: base + 'vision_model.onnx_data' }],
});
const { image_embeds } = await vision.run({
  pixel_values: new ort.Tensor('float32', pixelValues, [1, 3, 512, 512]),
});

The text model loads the same way with text_model.onnx / text_model.onnx_data and an input_ids tensor of shape [B, 64].

Limitations

  • The text model (565 MB in fp16) has to be downloaded and held in memory. If you only need image embeddings, load just the vision model.

License

Apache-2.0, as the original model. Original work by Google; see the SigLIP 2 paper and the original model card.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for upshift/siglip2-base-patch16-512-onnx

Quantized
(138)
this model

Collection including upshift/siglip2-base-patch16-512-onnx

Paper for upshift/siglip2-base-patch16-512-onnx