Video-Text-to-Text
Transformers
Safetensors
English
Chinese
qwen2_5_vl
image-text-to-text
video-understanding
multimodal
SWIM
Qwen2.5-VL
fine-grained-understanding
Eval Results (legacy)
text-generation-inference
Instructions to use BBBBCHAN/SWIM-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BBBBCHAN/SWIM-7B with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForImageTextToText processor = AutoProcessor.from_pretrained("BBBBCHAN/SWIM-7B") model = AutoModelForImageTextToText.from_pretrained("BBBBCHAN/SWIM-7B") - Notebooks
- Google Colab
- Kaggle
File size: 350 Bytes
a07fe85 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 | {
"min_pixels": 3136,
"max_pixels": 12845056,
"patch_size": 14,
"temporal_patch_size": 2,
"merge_size": 2,
"image_mean": [
0.48145466,
0.4578275,
0.40821073
],
"image_std": [
0.26862954,
0.26130258,
0.27577711
],
"image_processor_type": "Qwen2VLImageProcessor",
"processor_class": "Qwen2_5_VLProcessor"
} |