siglip2-base-patch16-512 — ONNX (browser-ready)
ONNX export of google/siglip2-base-patch16-512 (revision a89f5c5), split into a vision
and a text model, for onnxruntime-web on WebGPU and
wasm (and any other onnxruntime backend). fp32 and fp16, opset 18.
| File | Precision | Size | Inputs → output |
|---|---|---|---|
vision_model.onnx + vision_model.onnx_data |
fp32 | 374 MB | pixel_values float32 (B, 3, 512, 512) → image_embeds float32 (B, 768) |
text_model.onnx + text_model.onnx_data |
fp32 | 1.13 GB | input_ids int64 (B, 64) → text_embeds float32 (B, 768) |
vision_model_fp16.onnx + vision_model_fp16.onnx_data |
fp16 | 204 MB | same as fp32 |
text_model_fp16.onnx + text_model_fp16.onnx_data |
fp16 | 565 MB | same as fp32 |
B (batch) is dynamic; images are 512×512. Embeddings are not normalised.
config.json, preprocessor_config.json and the tokenizer files are copied from the original repo;
export_meta.json holds logit_scale (already exponentiated) and logit_bias for scoring.
fp16 variants (*_fp16.onnx) are for onnxruntime-web WebGPU (needs the shader-f16
feature): half the download, faster inference. Inputs and outputs stay float32, so they are drop-in
replacements. The vision model keeps its embeddings and attention-pooling head in fp32. WebGPU
computes in half precision, so expect small score shifts: in our tests image embeddings matched fp32 with
cosine ≥ 0.9995 and label probabilities moved by up to ~2 points. On wasm/CPU use the fp32 files.
Export notes
This is a fixed-resolution SigLIP2 checkpoint (SiglipModel architecture). The vision tower is exported
as-is for 512×512 inputs with a dynamic batch dimension; its learned position embeddings are tied to
that resolution.
Parity
fp32
Compared with the PyTorch model on onnxruntime (CPUExecutionProvider), threshold cosine ≥ 0.9999: passed.
| Tower | Inputs | Min cosine | Max abs diff |
|---|---|---|---|
| vision | 4 | 1.0000000 | 1.5e-05 |
| text | 6 | 1.0000000 | 1.8e-05 |
Max logit difference: 6.0e-05.
Versions: torch 2.14.1, transformers 5.19.0, onnxruntime 1.30.0.
fp16
Compared with the PyTorch model on onnxruntime (CPUExecutionProvider), threshold cosine ≥ 0.999: passed.
| Tower | Inputs | Min cosine | Max abs diff |
|---|---|---|---|
| vision | 4 | 0.9999950 | 9.4e-03 |
| text | 6 | 0.9999990 | 1.1e-02 |
Max logit difference: 2.9e-02.
Versions: torch 2.14.1, transformers 5.19.0, onnxruntime 1.30.0.
Inputs
Image. Resize to 512×512 (bilinear, antialiased when downscaling; the aspect ratio is not kept),
rescale to [-1, 1] (x / 127.5 - 1) and lay out channels-first as (3, 512, 512). Use
SiglipImageProcessor from transformers as the reference.
Text. Lowercase the text, tokenise with the included Gemma tokenizer.json, append <eos> and pad with
<pad> (id 0) to exactly 64 tokens. The model pools the last token, so the padding is required. In
Python, use Siglip2Tokenizer (transformers 5.x AutoTokenizer returns a GemmaTokenizer that does not
lowercase).
Scoring. p = sigmoid(logit_scale · cos(image, text) + logit_bias), per label. SigLIP scores labels
independently, so probabilities do not sum to 1; a correct label at 10–40 % with the rest near 0 is normal.
Usage with onnxruntime-web
import * as ort from 'onnxruntime-web/webgpu';
const base = 'https://huggingface.co/upshift/siglip2-base-patch16-512-onnx/resolve/main/';
const vision = await ort.InferenceSession.create(base + 'vision_model.onnx', {
executionProviders: ['webgpu', 'wasm'],
externalData: [{ path: 'vision_model.onnx_data', data: base + 'vision_model.onnx_data' }],
});
const { image_embeds } = await vision.run({
pixel_values: new ort.Tensor('float32', pixelValues, [1, 3, 512, 512]),
});
The text model loads the same way with text_model.onnx / text_model.onnx_data and an input_ids
tensor of shape [B, 64].
Limitations
- The text model (565 MB in fp16) has to be downloaded and held in memory. If you only need image embeddings, load just the vision model.
License
Apache-2.0, as the original model. Original work by Google; see the SigLIP 2 paper and the original model card.
- Downloads last month
- -
Model tree for upshift/siglip2-base-patch16-512-onnx
Base model
google/siglip2-base-patch16-512