Qwen3.5-9B for Janas-LLM
On a laptop running Debian 13 Linux, with an Intel Core Ultra 9 185H, its integrated Arc GPU and 32 GB of memory, Janas-LLM reads a prompt at 111 tok/s and writes at 14 tok/s, 24 tok/s with its own draft block in a chat: llama.cpp, on the same laptop, 78 and 8.7.
| On the same laptop | Janas-LLM | llama.cpp |
|---|---|---|
| reading a prompt of 1,024 tokens | 111 tok/s | 78 tok/s |
| writing (16 tokens) | 14 tok/s | 8.7 tok/s |
| writing in a chat, with its own draft block | 24 tok/s | not measured |
Measured on 9-11 October 2026: llama-bench -p 1024 -n 16 and janas-bench --prompt 1024 --gen 16, both at their best configuration; the drafts with janas-bench on a chat reply, the mean of three rounds (a draft guesses the next tokens and the model checks them: the same text, faster). On other machines the figures will differ. Details: Janas README.
The machine these figures come from
- Debian 13 (trixie) Linux, x86-64, kernel 6.12.107
- Intel Core Ultra 9 185H: 6 performance cores, 8 efficiency cores, 2 low-power cores; AVX2, FMA, F16C, AVX-VNNI
- its integrated Intel Arc GPU, through the Vulkan driver of Mesa 26.1.6
- 32 GB of DDR5-5600 memory
- a 1 TB NVMe SSD, about 5.7 GB/s reading
- llama.cpp build
2b18470(18 September 2026), Vulkan backend
These are the model's weights converted to JNS, the format of Janas-LLM, the inference engine in C of Janas, for ordinary computers, with or without a GPU. The files are read by Janas-LLM only.
Get it
Janas-LLM for Linux x86-64 (Ubuntu 22.04, Debian 12 and later), then the model:
curl -LO https://github.com/prabanta-dev/janas/releases/latest/download/janas-linux-x86_64.tar.gz
tar xf janas-linux-x86_64.tar.gz && cd janas
./janas-get qwen3.5-9b
./janas-chat qwen3.5-9b
janas-get downloads the files of this repository and checks them against the fingerprints below; janas-chat with no model shows the catalog and downloads the one you choose. Building Janas-LLM from its sources instead: prabanta-dev/janas.
Files
| File | Size | SHA-256 | Converted from |
|---|---|---|---|
qwen3.5-9b-q4km.jns |
5.68 GB | e40ec6c58ef2390c7a4ce8c9b7b6b384ed4e55c9980f646d1c98e83e2c018115 |
unsloth/Qwen3.5-9B-GGUF Qwen3.5-9B-Q4_K_M.gguf (catalog name qwen3.5-9b: Qwen3.5-9B, dense, with its MTP block) |
qwen3.5-9b-mtp-q4k.jns |
0.15 GB | 67f31f57dbc78ca8c02ef5e3aa137c5df3836a3decda20b6ea057d004089c361 |
the multi-token prediction tensors of Qwen/Qwen3.5-9B |
What was changed
These files are a modified work of the original model, as the Apache License 2.0 requires to be stated. Nothing was retrained or pruned: each file is a GGUF file quantized by the people named above, rewritten by Janas-LLM's gguf2jns in the order Janas reads it, with the metadata carried over. The multi-token prediction block, not in the GGUF, was taken from the original checkpoint and quantized to Q4_K by Janas-LLM's hf2jns_mtp.
Licence and attribution
Licensed under the Apache License 2.0, as the original model: see LICENSE.
- Model: Qwen/Qwen3.5-9B, by Alibaba Cloud (the Qwen team).
- Quantized GGUF: unsloth/Qwen3.5-9B-GGUF, by its authors.
- Converted to JNS by the Janas project with Janas-LLM (prabanta-dev/janas, GPL-3.0-or-later; the licence of the engine does not apply to these files).
LICENSE is the original repository's own.