YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

General conditions:

  • 32k context tokens.
  • Only GPU inference. No swaping/offloading.
  • Q8 KV cache quantization.
  • Flash Atention.
  • Evaluation Batch Size at 2048.
  • Physical Batch Size at 512.
  • MMAP and GPU KV cache offloading.
  • 1 concurrent requests.
  • Q3_K_XL
  • Nvidia RTX 4070 Laptop

MTP conditions:

  • Max draft tokens = 5
  • Min draft tokens = 0
  • Draft probability = 0.85

Platform:

  • ML Studio / CUDA 12 LLama.CPP

Results:

  • No quality tradeoff from base Alibaba's qwen 3.5 9B at Q4.

  • From 60 to 70 t/s depending on the code workload and the amount of context.

  • Enterprise speed feeling when used as harness model (claude code, open code and github copilot).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support