Granite-4.2-3B-NPU2
IBM Granite 4.2 3B quantised to the q4nx container that FastFlowLM loads, for AMD Ryzen AI NPUs (XDNA2 / NPU2).
The base model is IBM's, licensed Apache-2.0. This repository redistributes a quantised conversion of it under the same licence, with attribution. Nothing here is a new model; the weights are IBM's, in a different container.
Use
Requires a FastFlowLM build that includes the granite family. Support is not in an upstream release yet — see the pull request linked below.
flm run granite:3b
or place this directory as Granite-4.2-3B-NPU2 where FastFlowLM looks for
models.
What is in here
| file | |
|---|---|
model.q4nx |
the weights, Q4_1 semantics in the q4nx tile layout (2.6 GB) |
config.json |
granite geometry, plus a record of the folded multipliers |
tokenizer.json, tokenizer_config.json |
as shipped with the source package |
chat_template.jinja |
as shipped with the source package |
The tokenizer and chat template are deliberately unmodified. They look
mismatched — the tokenizer carries <|start_of_role|> as a single token while
the template emits ChatML markers that cost six tokens each — and replacing
either makes the model's output worse, which was measured rather than assumed:
| template | prompt | result |
|---|---|---|
| as shipped (ChatML) | 44 tokens | thinking trace, </think>, correct answer |
| role markers | 14 tokens | echoes the question first |
role markers + <think> |
15 tokens | URL-encoded garbage |
Swapping in upstream granite-4.2's own tokenizer is worse still: it and this one
differ on exactly 17 control-token ids out of 100352, so the model receives
<|pad|> where a turn marker should be, and the output degrades in a way that
stays fluent and is easy to misread as a model problem.
Geometry
Hidden 2560 · 40 layers · 40 query heads over 8 kv heads · head_dim 64 · intermediate 8192 · vocab 100352 · RoPE theta 1e7, half-split · RMSNorm eps 1e-5.
Granite needs head_dim 64 at hidden 2560, and every design FastFlowLM ships at hidden ≥ 2560 is head_dim 128, so it needed an engine of its own.
Provenance
Converted from IBM's published weights with the
q4nx-build tooling. The conversion
was verified tensor-by-tensor against the GGUF it was fed, and the resulting
forward pass was checked against an independent numpy implementation written
from config.json — cosine agreement at prefill and at decode steps 1 and 50.
Speed
FastFlowLM's granite engine currently runs on the CPU: about 8.7 tok/s decode on a Ryzen AI 9 HX 370. NPU kernels for this geometry exist and measure 13.6 tok/s of device time for the whole layer stack, but are not yet wired into the C++ dispatch path.
Links
- Base model: ibm-granite/granite-4.2-3b
- Runtime: ROCm/FastFlowLM
- Conversion tooling: Atomic-Germ/q4nx-build
- Downloads last month
- 27
Model tree for OpenFlowLM/Granite-4.2-3B-NPU2
Base model
ibm-granite/granite-4.1-3b-base