Video-Text-to-Text
Transformers
Safetensors
aviot_qwen
text-generation
video-language-model
video-question-answering
video-understanding
token-compression
optimal-transport
qwen2
siglip
Instructions to use ernie-research/AVIOT with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use ernie-research/AVIOT with Transformers:
# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("ernie-research/AVIOT", device_map="auto") - Notebooks
- Google Colab
- Kaggle
File size: 367 Bytes
a9ca494 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 | {
"do_normalize": true,
"do_rescale": true,
"do_resize": true,
"image_mean": [
0.5,
0.5,
0.5
],
"image_processor_type": "SiglipImageProcessor",
"image_std": [
0.5,
0.5,
0.5
],
"processor_class": "LlavaProcessor",
"resample": 3,
"rescale_factor": 0.00392156862745098,
"size": {
"height": 384,
"width": 384
}
}
|