Qwen3-4B for Janas-LLM
On a laptop running Debian 13 Linux, with an Intel Core Ultra 9 185H, its integrated Arc GPU and 32 GB of memory, Janas-LLM reads a prompt at 193 tok/s and writes at 27 tok/s: llama.cpp, on the same laptop, 164 and 18.
| On the same laptop | Janas-LLM | llama.cpp |
|---|---|---|
| reading a prompt of 1,024 tokens | 193 tok/s | 164 tok/s |
| writing (16 tokens) | 27 tok/s | 18 tok/s |
writing in a chat, Qwen3-0.6B drafting (--draft) |
26.3 tok/s (23.0 without, at the chat's temperature) | not measured |
The other files of this repository, measured the same way:
| File | reading: Janas-LLM | llama.cpp | writing: Janas-LLM | llama.cpp |
|---|---|---|---|---|
qwen3-4b-bq4km.jns (bartowski's Q4_K_M: same size, nearer the original) |
191 tok/s | 160 tok/s | 26 tok/s | 18 tok/s |
Measured on 9-11 October 2026: llama-bench -p 1024 -n 16 and janas-bench --prompt 1024 --gen 16, both at their best configuration; the draft pair in a chat reply at the chat's own temperature, as the README's section on --draft tells (a draft guesses the next tokens and the model checks them: the same text, faster). On other machines the figures will differ. Details: Janas README.
The machine these figures come from
- Debian 13 (trixie) Linux, x86-64, kernel 6.12.107
- Intel Core Ultra 9 185H: 6 performance cores, 8 efficiency cores, 2 low-power cores; AVX2, FMA, F16C, AVX-VNNI
- its integrated Intel Arc GPU, through the Vulkan driver of Mesa 26.1.6
- 32 GB of DDR5-5600 memory
- a 1 TB NVMe SSD, about 5.7 GB/s reading
- llama.cpp build
2b18470(18 September 2026), Vulkan backend
These are the model's weights converted to JNS, the format of Janas-LLM, the inference engine in C of Janas, for ordinary computers, with or without a GPU. The files are read by Janas-LLM only.
Get it
Janas-LLM for Linux x86-64 (Ubuntu 22.04, Debian 12 and later), then the model:
curl -LO https://github.com/prabanta-dev/janas/releases/latest/download/janas-linux-x86_64.tar.gz
tar xf janas-linux-x86_64.tar.gz && cd janas
./janas-get qwen3-4b # or: qwen3-4b-imatrix
./janas-chat qwen3-4b
janas-get downloads the files of this repository and checks them against the fingerprints below; janas-chat with no model shows the catalog and downloads the one you choose. Building Janas-LLM from its sources instead: prabanta-dev/janas.
Files
| File | Size | SHA-256 | Converted from |
|---|---|---|---|
qwen3-4b-q4km.jns |
2.50 GB | 1cbd8ebfdf14aee05777f5278e655382f040cdedc21d282e93806fc8cf279fc5 |
Qwen/Qwen3-4B-GGUF Qwen3-4B-Q4_K_M.gguf (catalog name qwen3-4b: Qwen3-4B, dense - the quick start) |
qwen3-4b-bq4km.jns |
2.50 GB | 134ae0afbbd3e45b3cdbfd36fd205fb98c91df62a129541048e29370a42b835e |
bartowski/Qwen_Qwen3-4B-GGUF Qwen_Qwen3-4B-Q4_K_M.gguf (catalog name qwen3-4b-imatrix: Qwen3-4B, bartowski's Q4_K_M: same size, nearer the original) |
What was changed
These files are a modified work of the original model, as the Apache License 2.0 requires to be stated. Nothing was retrained or pruned: each file is a GGUF file quantized by the people named above, rewritten by Janas-LLM's gguf2jns in the order Janas reads it, with the metadata carried over. The experts' down matrices were then cut by Janas-LLM's jns_planes into three planes of two bits each: a reordering of the same bits, which with all three planes gives back the quantized weights to the last bit.
Licence and attribution
Licensed under the Apache License 2.0, as the original model: see LICENSE.
- Model: Qwen/Qwen3-4B, by Alibaba Cloud (the Qwen team).
- Quantized GGUF: Qwen/Qwen3-4B-GGUF, by its authors.
- Quantized GGUF: bartowski/Qwen_Qwen3-4B-GGUF, by its authors.
- Converted to JNS by the Janas project with Janas-LLM (prabanta-dev/janas, GPL-3.0-or-later; the licence of the engine does not apply to these files).
LICENSE is the original repository's own.