Instructions to use nbeerbower/Qwen3.5-9B-Writing-DPO with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use nbeerbower/Qwen3.5-9B-Writing-DPO with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="nbeerbower/Qwen3.5-9B-Writing-DPO") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("nbeerbower/Qwen3.5-9B-Writing-DPO") model = AutoModelForCausalLM.from_pretrained("nbeerbower/Qwen3.5-9B-Writing-DPO", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use nbeerbower/Qwen3.5-9B-Writing-DPO with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "nbeerbower/Qwen3.5-9B-Writing-DPO" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nbeerbower/Qwen3.5-9B-Writing-DPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/nbeerbower/Qwen3.5-9B-Writing-DPO
- SGLang
How to use nbeerbower/Qwen3.5-9B-Writing-DPO with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "nbeerbower/Qwen3.5-9B-Writing-DPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nbeerbower/Qwen3.5-9B-Writing-DPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "nbeerbower/Qwen3.5-9B-Writing-DPO" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "nbeerbower/Qwen3.5-9B-Writing-DPO", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use nbeerbower/Qwen3.5-9B-Writing-DPO with Docker Model Runner:
docker model run hf.co/nbeerbower/Qwen3.5-9B-Writing-DPO
Qwen3.5 quantization sensitivity
Saw your model card says ORPO on Qwen3.5-9B with NF4 4-bit QLoRA. I’d heard Qwen3.5 was unusually sensitive to 4-bit finetuning/quantization error. Did you compare against bf16 LoRA at all, or notice instability during preference tuning?
Unfortunately do not have a bf16 lora, 4-bit quant with bitsandbytes is the only way to get everything to fit on my A6000. I'm sure it has some negative effect but honestly couldn't tell you how much. Didn't notice any obvious instability during the run, but I wasn't doing rigorous eval beyond loss curves and spot checks either.
You've piqued my curiosity though. I have hardware now to run a bf16 lora tune, so maybe I will try the same parameters as this and see how it turns out.
yes, qwen3.5 model family was reported to have bad results when trained on 4 bit, some people on reddit reported this and unsloth also suggested not using 4 bit training for these models as they are too sensitive to quantitisation when training, also i noticed you have been training coder models recently, you might also want to check out Jackrong’s reconstructed inverted traces dataset to create your own custom dpo dataset, feels like it could fit pretty well with the kind of preference-tuning work you’re doing.