siglip2-base-patch16-256 — ONNX (browser-ready)
ONNX export of google/siglip2-base-patch16-256 (revision 3f9f96c), split into a vision
and a text model, for onnxruntime-web on WebGPU and
wasm (and any other onnxruntime backend). fp32 and fp16, opset 18.
| File | Precision | Size | Inputs → output |
|---|---|---|---|
vision_model.onnx + vision_model.onnx_data |
fp32 | 372 MB | pixel_values float32 (B, 3, 256, 256) → image_embeds float32 (B, 768) |
text_model.onnx + text_model.onnx_data |
fp32 | 1.13 GB | input_ids int64 (B, 64) → text_embeds float32 (B, 768) |
vision_model_fp16.onnx + vision_model_fp16.onnx_data |
fp16 | 202 MB | same as fp32 |
text_model_fp16.onnx + text_model_fp16.onnx_data |
fp16 | 565 MB | same as fp32 |
B (batch) is dynamic; images are 256×256. Embeddings are not normalised.
config.json, preprocessor_config.json and the tokenizer files are copied from the original repo;
export_meta.json holds logit_scale (already exponentiated) and logit_bias for scoring.
fp16 variants (*_fp16.onnx) are for onnxruntime-web WebGPU (needs the shader-f16
feature): half the download, faster inference. Inputs and outputs stay float32, so they are drop-in
replacements. The vision model keeps its embeddings and attention-pooling head in fp32. WebGPU
computes in half precision, so expect small score shifts: in our tests image embeddings matched fp32 with
cosine ≥ 0.9995 and label probabilities moved by up to ~2 points. On wasm/CPU use the fp32 files.
Export notes
This is a fixed-resolution SigLIP2 checkpoint (SiglipModel architecture). The vision tower is exported
as-is for 256×256 inputs with a dynamic batch dimension; its learned position embeddings are tied to
that resolution.
Parity
fp32
Compared with the PyTorch model on onnxruntime (CPUExecutionProvider), threshold cosine ≥ 0.9999: passed.
| Tower | Inputs | Min cosine | Max abs diff |
|---|---|---|---|
| vision | 4 | 1.0000000 | 1.2e-05 |
| text | 6 | 1.0000000 | 1.3e-05 |
Max logit difference: 2.5e-05.
Versions: torch 2.14.1, transformers 5.19.0, onnxruntime 1.30.0.
fp16
Compared with the PyTorch model on onnxruntime (CPUExecutionProvider), threshold cosine ≥ 0.999: passed.
| Tower | Inputs | Min cosine | Max abs diff |
|---|---|---|---|
| vision | 4 | 0.9999987 | 3.5e-03 |
| text | 6 | 0.9999989 | 1.8e-02 |
Max logit difference: 1.6e-02.
Versions: torch 2.14.1, transformers 5.19.0, onnxruntime 1.30.0.
Inputs
Image. Resize to 256×256 (bilinear, antialiased when downscaling; the aspect ratio is not kept),
rescale to [-1, 1] (x / 127.5 - 1) and lay out channels-first as (3, 256, 256). Use
SiglipImageProcessor from transformers as the reference.
Text. Lowercase the text, tokenise with the included Gemma tokenizer.json, append <eos> and pad with
<pad> (id 0) to exactly 64 tokens. The model pools the last token, so the padding is required. In
Python, use Siglip2Tokenizer (transformers 5.x AutoTokenizer returns a GemmaTokenizer that does not
lowercase).
Scoring. p = sigmoid(logit_scale · cos(image, text) + logit_bias), per label. SigLIP scores labels
independently, so probabilities do not sum to 1; a correct label at 10–40 % with the rest near 0 is normal.
Usage with onnxruntime-web
import * as ort from 'onnxruntime-web/webgpu';
const base = 'https://huggingface.co/upshift/siglip2-base-patch16-256-onnx/resolve/main/';
const vision = await ort.InferenceSession.create(base + 'vision_model.onnx', {
executionProviders: ['webgpu', 'wasm'],
externalData: [{ path: 'vision_model.onnx_data', data: base + 'vision_model.onnx_data' }],
});
const { image_embeds } = await vision.run({
pixel_values: new ort.Tensor('float32', pixelValues, [1, 3, 256, 256]),
});
The text model loads the same way with text_model.onnx / text_model.onnx_data and an input_ids
tensor of shape [B, 64].
Limitations
- The text model (565 MB in fp16) has to be downloaded and held in memory. If you only need image embeddings, load just the vision model.
License
Apache-2.0, as the original model. Original work by Google; see the SigLIP 2 paper and the original model card.
- Downloads last month
- -
Model tree for upshift/siglip2-base-patch16-256-onnx
Base model
google/siglip2-base-patch16-256