Image-Text-to-Text
Transformers
Safetensors
qwen3_5
qwen
qwen3.5
multimodal
vision-language-model
conversational
Instructions to use Sfever/NekoQwen-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Sfever/NekoQwen-9B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="Sfever/NekoQwen-9B") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("Sfever/NekoQwen-9B") model = AutoModelForMultimodalLM.from_pretrained("Sfever/NekoQwen-9B", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Sfever/NekoQwen-9B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Sfever/NekoQwen-9B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sfever/NekoQwen-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/Sfever/NekoQwen-9B
- SGLang
How to use Sfever/NekoQwen-9B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Sfever/NekoQwen-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sfever/NekoQwen-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Sfever/NekoQwen-9B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Sfever/NekoQwen-9B", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use Sfever/NekoQwen-9B with Docker Model Runner:
docker model run hf.co/Sfever/NekoQwen-9B
NekoQwen-9B
Qwen3.5-9B finetuned by NekoQA-30K
Model Details
- Architecture:
Qwen3_5ForConditionalGeneration - Processor:
Qwen3VLProcessor - Precision:
float16 - Format: sharded
safetensors - Parameter count: about 9.41B
- Repository size: about 18 GB
- Modalities: text, image, and video inputs with text generation output
- Max position embeddings:
262144 - Transformers version in config:
5.3.0
Fine-Tuning Summary
- Base model:
Qwen/Qwen3.5-9B - Tuning method: LoRA merged into full weights
- Epochs:
1.0 - Learning rate:
1e-4 - Per-device batch size:
1 - Gradient accumulation:
16 - Sequence length:
768 - Precision during training:
fp16
Usage
import torch
from transformers import AutoProcessor, Qwen3_5ForConditionalGeneration
model_id = "your-username/your-repo"
processor = AutoProcessor.from_pretrained(model_id)
model = Qwen3_5ForConditionalGeneration.from_pretrained(
model_id,
torch_dtype=torch.float16,
device_map="auto",
)
messages = [
{
"role": "user",
"content": [
{"type": "text", "text": "Describe the main characteristics of this model in one paragraph."},
],
}
]
text = processor.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
inputs = processor(text=[text], padding=True, return_tensors="pt").to(model.device)
generated_ids = model.generate(**inputs, max_new_tokens=128)
print(processor.batch_decode(generated_ids, skip_special_tokens=True)[0])
For image or video inputs, use the same chat-template message structure with Qwen3VLProcessor.
Notes
This folder contains the merged checkpoint, tokenizer, processor configuration, and chat template needed to load the model with Transformers.
Training data provenance, evaluation results, and intended-use notes are not documented in this folder yet. Add those details before making the repository public if you want a complete public model card.
- Downloads last month
- 5