Qwen3-Coder-Next for Janas-LLM
On a laptop running Debian 13 Linux, with an Intel Core Ultra 9 185H, its integrated Arc GPU and 32 GB of memory, Janas-LLM reads a prompt at 66 tok/s and writes at 22 tok/s: llama.cpp, on the same laptop, 14 and 2.2.
| On the same laptop | Janas-LLM | llama.cpp |
|---|---|---|
| reading a prompt of 1,024 tokens | 66 tok/s | 14 tok/s |
| writing (16 tokens) | 22 tok/s | 2.2 tok/s |
Measured on 9-11 October 2026: llama-bench -p 1024 -n 16 and janas-bench --prompt 1024 --gen 16, both at their best configuration. On other machines the figures will differ. Details: Janas README.
The machine these figures come from
- Debian 13 (trixie) Linux, x86-64, kernel 6.12.107
- Intel Core Ultra 9 185H: 6 performance cores, 8 efficiency cores, 2 low-power cores; AVX2, FMA, F16C, AVX-VNNI
- its integrated Intel Arc GPU, through the Vulkan driver of Mesa 26.1.6
- 32 GB of DDR5-5600 memory
- a 1 TB NVMe SSD, about 5.7 GB/s reading
- llama.cpp build
2b18470(18 September 2026), Vulkan backend
The model (48.4 GB) is larger than the laptop's memory: Janas-LLM streams from the NVMe SSD the experts each token needs and keeps the ones that keep coming back in memory. llama.cpp was measured as on Qwen3-Next-80B, with a build without weight repacking: reading with the GPU and the experts left in the file (-ncmoe 48 --no-host 1 -lm mmap), writing on the CPU alone, the faster of its ways for each.
These are the model's weights converted to JNS, the format of Janas-LLM, the inference engine in C of Janas, for ordinary computers, with or without a GPU. The files are read by Janas-LLM only.
Get it
Janas-LLM for Linux x86-64 (Ubuntu 22.04, Debian 12 and later), then the model:
curl -LO https://github.com/prabanta-dev/janas/releases/latest/download/janas-linux-x86_64.tar.gz
tar xf janas-linux-x86_64.tar.gz && cd janas
./janas-get qwen3-coder-next
./janas-chat qwen3-coder-next
janas-get downloads the files of this repository and checks them against the fingerprints below; janas-chat with no model shows the catalog and downloads the one you choose. Building Janas-LLM from its sources instead: prabanta-dev/janas.
Files
| File | Size | SHA-256 | Converted from |
|---|---|---|---|
qwen3-coder-next-q4km.jns |
48.41 GB | 0276b0697a05075fec8ab076a757205537804ff41dd7e9dadace7c6993763563 |
Qwen/Qwen3-Coder-Next-GGUF Qwen3-Coder-Next-Q4_K_M/Qwen3-Coder-Next-Q4_K_M-00001-of-00004.gguf (catalog name qwen3-coder-next: Qwen3-Coder-Next, mixture of experts, for code) |
What was changed
These files are a modified work of the original model, as the Apache License 2.0 requires to be stated. Nothing was retrained or pruned: each file is a GGUF file quantized by the people named above, rewritten by Janas-LLM's gguf2jns in the order Janas reads it, with the metadata carried over. Each (layer, expert) has a slot of its own, so that the experts a token needs can be read from disk on their own. The experts' down matrices were then cut by Janas-LLM's jns_planes into three planes of two bits each: a reordering of the same bits, which with all three planes gives back the quantized weights to the last bit.
Licence and attribution
Licensed under the Apache License 2.0, as the original model: see LICENSE.
- Model: Qwen/Qwen3-Coder-Next, by Alibaba Cloud (the Qwen team).
- Quantized GGUF: Qwen/Qwen3-Coder-Next-GGUF, by its authors.
- Converted to JNS by the Janas project with Janas-LLM (prabanta-dev/janas, GPL-3.0-or-later; the licence of the engine does not apply to these files).
The original repository carries no LICENSE file; its card declares Apache 2.0. LICENSE here is the licence's text as the Apache Software Foundation publishes it.
Model tree for prabanta-dev/Qwen3-Coder-Next-JNS
Base model
Qwen/Qwen3-Coder-Next