|
Download README.md from Ambarella/SigLIP2: direct link, hf CLI and curl.
- Browser
- Download file 2.95 kB
-
https://huggingface.co/Ambarella/SigLIP2/resolve/main/README.md
- Command line
-
hf download hf://Ambarella/SigLIP2/README.md
-
curl -L -o README.md https://huggingface.co/Ambarella/SigLIP2/resolve/main/README.md
2.95 kB
metadata
library_name: pytorch
SigLIP 2 is a multilingual vision–language encoder that extends the original SigLIP training objective with captioning, self-distillation, masked prediction, and improved data curation for stronger semantic understanding, localization, and dense visual representations.
SigLIP2-Base-Patch16-224
This model uses the SigLIP 2 Base-Patch16-224x224 variant, based on a ViT-Base vision encoder with 16×16 image patches and approximately 86M parameters. It is well suited for zero-shot image classification, image–text retrieval, and as a vision encoder for VLMs and downstream vision tasks.
Model Configuration:
- Reference implementation: google/siglip2-base-patch16-224
- Original Weight: SigLIP2-Base-Patch16
- Dataset: ImageNet
- Resolution: 3x224x224
- Support Cooper version:
- Cooper SDK: [2.5.4]
- Cooper Foundry: [2.3]
| Model | Device | Model Link |
|---|---|---|
| SigLIP2-Base-Patch16 Image Encoder | N1-655 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | N1-655 | Model_Link |
| SigLIP2-Base-Patch16 Image Encoder | X7 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | X7 | Model_Link |
| SigLIP2-Base-Patch16 Image Encoder | CV7 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | CV7 | Model_Link |
| SigLIP2-Base-Patch16 Image Encoder | CV72 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | CV72 | Model_Link |
| SigLIP2-Base-Patch16 Image Encoder | CV75 | Model_Link |
| SigLIP2-Base-Patch16 Text Encoder | CV75 | Model_Link |
