HanzoHuang commited on
Commit
90dedc0
·
verified ·
1 Parent(s): 8a667d2

Add complete model card

Browse files
Files changed (1) hide show
  1. README.md +59 -0
README.md CHANGED
@@ -1,3 +1,62 @@
1
  ---
2
  license: apache-2.0
 
 
 
 
 
 
 
 
 
 
 
3
  ---
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
  ---
2
  license: apache-2.0
3
+ base_model: Qwen/Qwen3.5-4B
4
+ pipeline_tag: image-text-to-text
5
+ library_name: rkllm
6
+ tags:
7
+ - rkllm
8
+ - rockchip
9
+ - rk3576
10
+ - rk3588
11
+ - qwen
12
+ - qwen3.5
13
+ - multimodal
14
  ---
15
+
16
+ # Qwen3.5-4B-RKLLM
17
+
18
+ RKLLM-converted Qwen3.5-4B artifacts for deployment on Rockchip RK3576 and RK3588 NPUs. This repository includes language-model binaries and the matching vision encoder artifacts for multimodal inference.
19
+
20
+ `.rkllm` and `.rknn` files require the Rockchip RKLLM/RKNN runtime. They cannot be loaded directly with Transformers, llama.cpp, or Ollama.
21
+
22
+ ## Base Model
23
+
24
+ - Model: Qwen3.5-4B
25
+ - Author: Qwen Team
26
+ - Original model: [Qwen/Qwen3.5-4B](https://huggingface.co/Qwen/Qwen3.5-4B)
27
+ - Original license: Apache-2.0
28
+
29
+ Refer to the upstream model card for the original model's capabilities, limitations, and acceptable-use guidance.
30
+
31
+ ## Available Artifacts
32
+
33
+ | Target | Language model | Quantization | Vision model |
34
+ | --- | --- | --- | --- |
35
+ | RK3576 | [`Qwen3.5-4B_RK3576_w4a16_g128.rkllm`](RK3576/Qwen3.5-4B_RK3576_w4a16_g128.rkllm) | W4A16 (g128) | [`Qwen3.5-4B_vision_RK3576.rknn`](RK3576/Qwen3.5-4B_vision_RK3576.rknn) |
36
+ | RK3576 | [`Qwen3.5-4B_RK3576_w8a8.rkllm`](RK3576/Qwen3.5-4B_RK3576_w8a8.rkllm) | W8A8 | [`Qwen3.5-4B_vision_RK3576.rknn`](RK3576/Qwen3.5-4B_vision_RK3576.rknn) |
37
+ | RK3588 | [`Qwen3.5-4B_RK3588_w8a8.rkllm`](RK3588/Qwen3.5-4B_RK3588_w8a8.rkllm) | W8A8 | [`Qwen3.5-4B_vision_RK3588.rknn`](RK3588/Qwen3.5-4B_vision_RK3588.rknn) |
38
+
39
+ The repository also contains `Qwen3.5-4B_vision.onnx`, the vision encoder in ONNX format, and `Qwen3.5-4B_data_quant.json`, the calibration data used during conversion.
40
+
41
+ ## Usage
42
+
43
+ Download one language-model artifact and the matching vision artifact for your SoC. Do not mix RK3576 and RK3588 files.
44
+
45
+ ```bash
46
+ hf download HanzoHuang/Qwen3.5-4B-RKLLM \
47
+ RK3576/Qwen3.5-4B_RK3576_w4a16_g128.rkllm \
48
+ RK3576/Qwen3.5-4B_vision_RK3576.rknn \
49
+ --local-dir Qwen3.5-4B-RKLLM
50
+ ```
51
+
52
+ Load the `.rkllm` model with a compatible RKLLM runtime and the `.rknn` vision encoder with the matching RKNN runtime. Application code must implement Qwen3.5's multimodal preprocessing and prompt format.
53
+
54
+ ## Limitations
55
+
56
+ - These are hardware-specific converted artifacts, not Transformers checkpoints.
57
+ - Runtime, driver, and toolkit compatibility depends on the Rockchip software stack installed on the device.
58
+ - Conversion may change output quality relative to the upstream floating-point model; validate on your own workload.
59
+
60
+ ## Acknowledgements
61
+
62
+ Thanks to the Qwen Team for releasing Qwen3.5-4B and to Rockchip and RKLLM contributors for the deployment toolchain.