Instructions to use mconcat/Qwopus3.5-27B-v3-FP8-Dynamic with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use mconcat/Qwopus3.5-27B-v3-FP8-Dynamic with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="mconcat/Qwopus3.5-27B-v3-FP8-Dynamic") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("mconcat/Qwopus3.5-27B-v3-FP8-Dynamic") model = AutoModelForMultimodalLM.from_pretrained("mconcat/Qwopus3.5-27B-v3-FP8-Dynamic", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use mconcat/Qwopus3.5-27B-v3-FP8-Dynamic with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "mconcat/Qwopus3.5-27B-v3-FP8-Dynamic" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mconcat/Qwopus3.5-27B-v3-FP8-Dynamic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/mconcat/Qwopus3.5-27B-v3-FP8-Dynamic
- SGLang
How to use mconcat/Qwopus3.5-27B-v3-FP8-Dynamic with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "mconcat/Qwopus3.5-27B-v3-FP8-Dynamic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mconcat/Qwopus3.5-27B-v3-FP8-Dynamic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "mconcat/Qwopus3.5-27B-v3-FP8-Dynamic" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "mconcat/Qwopus3.5-27B-v3-FP8-Dynamic", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use mconcat/Qwopus3.5-27B-v3-FP8-Dynamic with Docker Model Runner:
docker model run hf.co/mconcat/Qwopus3.5-27B-v3-FP8-Dynamic
My recipe for deployment
Using vllm, u need to manually install transformers>=5.3 to fit new RoPE embedding.
My hardware: RTX3090 + RTX4090
deployment script:
# Enable memory profiler to estimate CUDA graphs v0.19 functionality
export VLLM_MEMORY_PROFILER_ESTIMATE_CUDAGRAPHS=1
export MODEL_NAME="mconcat/Qwopus3.5-27B-v3-FP8-Dynamic"
# Start vLLM with reduced swap space
vllm serve $MODEL_NAME \
--served-model-name vllm/Qwen3.5-27B \
--trust-remote-code \
--tensor-parallel-size 2 \
--max-model-len 219520 \
--gpu-memory-utilization 0.92 \
--enable-auto-tool-choice \
--enable-chunked-prefill \
--enable-prefix-caching \
--max-num-batched-tokens 4096 \
--max-num-seqs 4 \
--kv-cache-dtype fp8 \
--tool-call-parser hermes \
--reasoning-parser qwen3 \
--no-use-tqdm-on-load \
--host 0.0.0.0 \
--port 8000 \
--language-model-only
I feel the stability is a bit improved, the original 27B fail and melfunctioned on tool calling, but this one is fine so far. However i havent try the long context conversation / agentic coding yet. Any one have data on it?
Thanks for sharing!
Quick question β why use --language-model-only? This model supports vision (image-text-to-text).
I use the Claude code. After running a few rounds, it would suddenly stop and then say "Continue" before resuming the run.
Thanks for sharing!
Quick question β why use --language-model-only? This model supports vision (image-text-to-text).
Simply ofcaz i dont need the image part, and it can save some RAM by stopping it
I use the Claude code. After running a few rounds, it would suddenly stop and then say "Continue" before resuming the run.
Do you means Claude code will automatically type "Continue" in the text box and let it run? Or you have to manually type "Continue" to let this model run ?
