Instructions to use Qwen/Qwen-VL-Chat with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Qwen/Qwen-VL-Chat with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Qwen/Qwen-VL-Chat", trust_remote_code=True, device_map="auto")# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("Qwen/Qwen-VL-Chat", trust_remote_code=True, dtype="auto", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Qwen/Qwen-VL-Chat with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Qwen/Qwen-VL-Chat" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen-VL-Chat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/Qwen/Qwen-VL-Chat
- SGLang
How to use Qwen/Qwen-VL-Chat with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Qwen/Qwen-VL-Chat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen-VL-Chat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Qwen/Qwen-VL-Chat" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Qwen/Qwen-VL-Chat", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use Qwen/Qwen-VL-Chat with Docker Model Runner:
docker model run hf.co/Qwen/Qwen-VL-Chat
Unofficial LLaMAfied Version in HF format - 非官方的LLaMA化HF格式版本
https://huggingface.co/JosephusCheung/Qwen-VL-LLaMAfied-7B-Chat
Similar to LLaMAfied Qwen-7B-Chat, the visual parts and LLM are separated, and LLM is restructured and recalibrated into the standard LLaMA/LLaMA2 format (with GPT-2 tokenizer). It can be used with any tools that are compatible with LLaMA, such as stream output, llama.cpp quantization, and so on.
Thanks for your work. I think the vision.bin file is the visual parts, How to fine tuning this model and inference?
Thanks for your work. I think the vision.bin file is the visual parts, How to fine tuning this model and inference?
The structure of the LLM part is identical to that of LLaMA, allowing you to utilize HF transformers with the LlamaForCausalLM. For the vision part, you can utilize visual.py from the original Qwen-VL repository. This allows you to convert images into LM input embeddings. You can then manually concatenate these with your text instruction input for the LLM.
It is quite obvious I think.
https://huggingface.co/JosephusCheung/Qwen-VL-LLaMAfied-7B-Chat
Similar to LLaMAfied Qwen-7B-Chat, the visual parts and LLM are separated, and LLM is restructured and recalibrated into the standard LLaMA/LLaMA2 format (with GPT-2 tokenizer). It can be used with any tools that are compatible with LLaMA, such as stream output, llama.cpp quantization, and so on.
how to use?
https://huggingface.co/JosephusCheung/Qwen-VL-LLaMAfied-7B-Chat
Similar to LLaMAfied Qwen-7B-Chat, the visual parts and LLM are separated, and LLM is restructured and recalibrated into the standard LLaMA/LLaMA2 format (with GPT-2 tokenizer). It can be used with any tools that are compatible with LLaMA, such as stream output, llama.cpp quantization, and so on.
how to use?
Use LLM the way you use LLaMA-2, and use visual.py from Qwen-VL for VL part, which is obvious for literate people.
https://huggingface.co/JosephusCheung/Qwen-VL-LLaMAfied-7B-Chat
Similar to LLaMAfied Qwen-7B-Chat, the visual parts and LLM are separated, and LLM is restructured and recalibrated into the standard LLaMA/LLaMA2 format (with GPT-2 tokenizer). It can be used with any tools that are compatible with LLaMA, such as stream output, llama.cpp quantization, and so on.
how to use?
Use LLM the way you use LLaMA-2, and use visual.py from Qwen-VL for VL part, which is obvious for literate people.
This is difficult for me, i don't know how to merge two models for inferencing. could you give me some code examples? Thanks!
https://huggingface.co/JosephusCheung/Qwen-VL-LLaMAfied-7B-Chat
Similar to LLaMAfied Qwen-7B-Chat, the visual parts and LLM are separated, and LLM is restructured and recalibrated into the standard LLaMA/LLaMA2 format (with GPT-2 tokenizer). It can be used with any tools that are compatible with LLaMA, such as stream output, llama.cpp quantization, and so on.
Great work, thank you very much.