Feature Extraction
Transformers
Safetensors
avito_gated_fusion
siglip
siglip2
vision
text
clip
multimodal
image-text-embeddings
pet-recognition
custom_code
Instructions to use AvitoTech/SigLIP2-giant-e5small-v2-gating with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AvitoTech/SigLIP2-giant-e5small-v2-gating with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("feature-extraction", model="AvitoTech/SigLIP2-giant-e5small-v2-gating", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("AvitoTech/SigLIP2-giant-e5small-v2-gating", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 1,701 Bytes
c98d3b8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 | {
"model_type": "avito_gated_fusion",
"architectures": [
"AvitoGatedFusionModel"
],
"auto_map": {
"AutoConfig": "configuration_avito_gated.AvitoGatedFusionConfig",
"AutoModel": "modeling_avito_gated.AvitoGatedFusionModel"
},
"embedding_dim": 512,
"gate_hidden_dim": 128,
"siglip_config": {
"initializer_factor": 1.0,
"model_type": "siglip",
"text_config": {
"vocab_size": 256000,
"hidden_size": 1152,
"intermediate_size": 4304,
"num_hidden_layers": 27,
"num_attention_heads": 16,
"max_position_embeddings": 64,
"layer_norm_eps": 1e-06,
"hidden_act": "gelu_pytorch_tanh",
"attention_dropout": 0.0,
"projection_size": 1536,
"model_type": "siglip_text_model"
},
"vision_config": {
"hidden_size": 1536,
"intermediate_size": 6144,
"num_hidden_layers": 40,
"num_attention_heads": 16,
"num_channels": 3,
"patch_size": 16,
"image_size": 384,
"attention_dropout": 0.0,
"layer_norm_eps": 1e-06,
"hidden_act": "gelu_pytorch_tanh",
"model_type": "siglip_vision_model"
}
},
"text_config": {
"pad_token_id": 0,
"model_type": "bert",
"vocab_size": 30522,
"hidden_size": 384,
"num_hidden_layers": 12,
"num_attention_heads": 12,
"hidden_act": "gelu",
"intermediate_size": 1536,
"hidden_dropout_prob": 0.1,
"attention_probs_dropout_prob": 0.1,
"max_position_embeddings": 512,
"type_vocab_size": 2,
"initializer_range": 0.02,
"layer_norm_eps": 1e-12,
"position_embedding_type": "absolute",
"use_cache": true,
"classifier_dropout": null
},
"dtype": "float32"
}
|