Visual Question Answering
Transformers
Safetensors
English
videorefer_qwen2
text-generation
multimodal large language model
large video-language model
Instructions to use DAMO-NLP-SG/VideoRefer-7B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DAMO-NLP-SG/VideoRefer-7B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("visual-question-answering", model="DAMO-NLP-SG/VideoRefer-7B")# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("DAMO-NLP-SG/VideoRefer-7B", dtype="auto") - Notebooks
- Google Colab
- Kaggle
Any tutorial ipynb for fine-tuning the model?
#1
by chenxiangyi10 - opened
Thanks for sharing your great work. It would be helpful if a tutorial on fine-tuning the model could be released. This would definitely increase practitioners' engagement.