Qwen3.5-9B for Janas-LLM

On a laptop running Debian 13 Linux, with an Intel Core Ultra 9 185H, its integrated Arc GPU and 32 GB of memory, Janas-LLM reads a prompt at 111 tok/s and writes at 14 tok/s, 24 tok/s with its own draft block in a chat: llama.cpp, on the same laptop, 78 and 8.7.

On the same laptop Janas-LLM llama.cpp
reading a prompt of 1,024 tokens 111 tok/s 78 tok/s
writing (16 tokens) 14 tok/s 8.7 tok/s
writing in a chat, with its own draft block 24 tok/s not measured

Measured on 9-11 October 2026: llama-bench -p 1024 -n 16 and janas-bench --prompt 1024 --gen 16, both at their best configuration; the drafts with janas-bench on a chat reply, the mean of three rounds (a draft guesses the next tokens and the model checks them: the same text, faster). On other machines the figures will differ. Details: Janas README.

The machine these figures come from

  • Debian 13 (trixie) Linux, x86-64, kernel 6.12.107
  • Intel Core Ultra 9 185H: 6 performance cores, 8 efficiency cores, 2 low-power cores; AVX2, FMA, F16C, AVX-VNNI
  • its integrated Intel Arc GPU, through the Vulkan driver of Mesa 26.1.6
  • 32 GB of DDR5-5600 memory
  • a 1 TB NVMe SSD, about 5.7 GB/s reading
  • llama.cpp build 2b18470 (18 September 2026), Vulkan backend

These are the model's weights converted to JNS, the format of Janas-LLM, the inference engine in C of Janas, for ordinary computers, with or without a GPU. The files are read by Janas-LLM only.

Get it

Janas-LLM for Linux x86-64 (Ubuntu 22.04, Debian 12 and later), then the model:

curl -LO https://github.com/prabanta-dev/janas/releases/latest/download/janas-linux-x86_64.tar.gz
tar xf janas-linux-x86_64.tar.gz && cd janas
./janas-get qwen3.5-9b
./janas-chat qwen3.5-9b

janas-get downloads the files of this repository and checks them against the fingerprints below; janas-chat with no model shows the catalog and downloads the one you choose. Building Janas-LLM from its sources instead: prabanta-dev/janas.

Files

File Size SHA-256 Converted from
qwen3.5-9b-q4km.jns 5.68 GB e40ec6c58ef2390c7a4ce8c9b7b6b384ed4e55c9980f646d1c98e83e2c018115 unsloth/Qwen3.5-9B-GGUF Qwen3.5-9B-Q4_K_M.gguf (catalog name qwen3.5-9b: Qwen3.5-9B, dense, with its MTP block)
qwen3.5-9b-mtp-q4k.jns 0.15 GB 67f31f57dbc78ca8c02ef5e3aa137c5df3836a3decda20b6ea057d004089c361 the multi-token prediction tensors of Qwen/Qwen3.5-9B

What was changed

These files are a modified work of the original model, as the Apache License 2.0 requires to be stated. Nothing was retrained or pruned: each file is a GGUF file quantized by the people named above, rewritten by Janas-LLM's gguf2jns in the order Janas reads it, with the metadata carried over. The multi-token prediction block, not in the GGUF, was taken from the original checkpoint and quantized to Q4_K by Janas-LLM's hf2jns_mtp.

Licence and attribution

Licensed under the Apache License 2.0, as the original model: see LICENSE.

LICENSE is the original repository's own.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prabanta-dev/Qwen3.5-9B-JNS

Finetuned
Qwen/Qwen3.5-9B
Quantized
(567)
this model

Collection including prabanta-dev/Qwen3.5-9B-JNS