Image Feature Extraction
Transformers
Safetensors
English
videollama3_vision_encoder
feature-extraction
visual-encoder
multi-modal-large-language-model
custom_code
Instructions to use DAMO-NLP-SG/VL3-SigLIP-NaViT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DAMO-NLP-SG/VL3-SigLIP-NaViT with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-feature-extraction", model="DAMO-NLP-SG/VL3-SigLIP-NaViT", trust_remote_code=True)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("DAMO-NLP-SG/VL3-SigLIP-NaViT", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
Does this only supports image?
#6
by 2U1 - opened
I'm not sure because that the model has 2 version Image and video.
Does this vision encoder supports video?
If so, can I get some example how can I get the embedding from videos.
Also, Can I get some examples for multi-image too?