Qwen3-4B for Janas-LLM

On a laptop running Debian 13 Linux, with an Intel Core Ultra 9 185H, its integrated Arc GPU and 32 GB of memory, Janas-LLM reads a prompt at 193 tok/s and writes at 27 tok/s: llama.cpp, on the same laptop, 164 and 18.

On the same laptop Janas-LLM llama.cpp
reading a prompt of 1,024 tokens 193 tok/s 164 tok/s
writing (16 tokens) 27 tok/s 18 tok/s
writing in a chat, Qwen3-0.6B drafting (--draft) 26.3 tok/s (23.0 without, at the chat's temperature) not measured

The other files of this repository, measured the same way:

File reading: Janas-LLM llama.cpp writing: Janas-LLM llama.cpp
qwen3-4b-bq4km.jns (bartowski's Q4_K_M: same size, nearer the original) 191 tok/s 160 tok/s 26 tok/s 18 tok/s

Measured on 9-11 October 2026: llama-bench -p 1024 -n 16 and janas-bench --prompt 1024 --gen 16, both at their best configuration; the draft pair in a chat reply at the chat's own temperature, as the README's section on --draft tells (a draft guesses the next tokens and the model checks them: the same text, faster). On other machines the figures will differ. Details: Janas README.

The machine these figures come from

  • Debian 13 (trixie) Linux, x86-64, kernel 6.12.107
  • Intel Core Ultra 9 185H: 6 performance cores, 8 efficiency cores, 2 low-power cores; AVX2, FMA, F16C, AVX-VNNI
  • its integrated Intel Arc GPU, through the Vulkan driver of Mesa 26.1.6
  • 32 GB of DDR5-5600 memory
  • a 1 TB NVMe SSD, about 5.7 GB/s reading
  • llama.cpp build 2b18470 (18 September 2026), Vulkan backend

These are the model's weights converted to JNS, the format of Janas-LLM, the inference engine in C of Janas, for ordinary computers, with or without a GPU. The files are read by Janas-LLM only.

Get it

Janas-LLM for Linux x86-64 (Ubuntu 22.04, Debian 12 and later), then the model:

curl -LO https://github.com/prabanta-dev/janas/releases/latest/download/janas-linux-x86_64.tar.gz
tar xf janas-linux-x86_64.tar.gz && cd janas
./janas-get qwen3-4b                # or: qwen3-4b-imatrix
./janas-chat qwen3-4b

janas-get downloads the files of this repository and checks them against the fingerprints below; janas-chat with no model shows the catalog and downloads the one you choose. Building Janas-LLM from its sources instead: prabanta-dev/janas.

Files

File Size SHA-256 Converted from
qwen3-4b-q4km.jns 2.50 GB 1cbd8ebfdf14aee05777f5278e655382f040cdedc21d282e93806fc8cf279fc5 Qwen/Qwen3-4B-GGUF Qwen3-4B-Q4_K_M.gguf (catalog name qwen3-4b: Qwen3-4B, dense - the quick start)
qwen3-4b-bq4km.jns 2.50 GB 134ae0afbbd3e45b3cdbfd36fd205fb98c91df62a129541048e29370a42b835e bartowski/Qwen_Qwen3-4B-GGUF Qwen_Qwen3-4B-Q4_K_M.gguf (catalog name qwen3-4b-imatrix: Qwen3-4B, bartowski's Q4_K_M: same size, nearer the original)

What was changed

These files are a modified work of the original model, as the Apache License 2.0 requires to be stated. Nothing was retrained or pruned: each file is a GGUF file quantized by the people named above, rewritten by Janas-LLM's gguf2jns in the order Janas reads it, with the metadata carried over. The experts' down matrices were then cut by Janas-LLM's jns_planes into three planes of two bits each: a reordering of the same bits, which with all three planes gives back the quantized weights to the last bit.

Licence and attribution

Licensed under the Apache License 2.0, as the original model: see LICENSE.

LICENSE is the original repository's own.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prabanta-dev/Qwen3-4B-JNS

Finetuned
Qwen/Qwen3-4B
Quantized
(343)
this model

Collection including prabanta-dev/Qwen3-4B-JNS