YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
General conditions:
- 32k context tokens.
- Only GPU inference. No swaping/offloading.
- Q8 KV cache quantization.
- Flash Atention.
- Evaluation Batch Size at 2048.
- Physical Batch Size at 512.
- MMAP and GPU KV cache offloading.
- 1 concurrent requests.
- Q3_K_XL
- Nvidia RTX 4070 Laptop
MTP conditions:
- Max draft tokens = 5
- Min draft tokens = 0
- Draft probability = 0.85
Platform:
- ML Studio / CUDA 12 LLama.CPP
Results:
No quality tradeoff from base Alibaba's qwen 3.5 9B at Q4.
From 60 to 70 t/s depending on the code workload and the amount of context.
Enterprise speed feeling when used as harness model (claude code, open code and github copilot).
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support