--- license: apache-2.0 base_model: Qwen/Qwen3-VL-4B-Instruct tags: - rkllm - rknn - rk3588 - rockchip - npu - quantized - vision-language - multimodal - qwen3 language: - en - zh pipeline_tag: image-text-to-text --- # Qwen3-VL-4B-Instruct — RKLLM v1.2.3 (w8a8, RK3588) RKLLM/RKNN conversion of [Qwen/Qwen3-VL-4B-Instruct](https://huggingface.co/Qwen/Qwen3-VL-4B-Instruct) for Rockchip RK3588 NPU inference. Converted with RKLLM Toolkit v1.2.3 (language model) and RKNN Toolkit (vision encoder). This is a multimodal vision-language model — it accepts both images and text as input. ## Key Details | Property | Value | |----------|-------| | Base Model | Qwen/Qwen3-VL-4B-Instruct | | Toolkit Version | RKLLM Toolkit v1.2.3 / RKNN Toolkit | | Runtime Version | RKLLM Runtime ≥ v1.2.1 + RKNN Runtime | | Quantization | w8a8 (8-bit weights, 8-bit activations) | | Target Platform | RK3588 | | NPU Cores | 3 | | Thinking Mode | ❌ Disabled | | Model Type | Vision-Language (VLM) | | Languages | English, Chinese (multilingual) | ## Why This Model? Qwen3-VL-4B-Instruct is Alibaba's 4B vision-language model. It handles image understanding, visual QA, document analysis, and chart reading with strong multilingual support. Running on the RK3588 NPU enables fully local, GPU-free multimodal inference. Compared to the smaller Qwen3-VL-2B, the 4B variant offers meaningfully better image understanding and text extraction. ## Hardware Tested - **Orange Pi 5 Plus** — RK3588, 16GB RAM, Armbian Linux - RKNPU driver 0.9.8 - RKLLM Runtime v1.2.3 ## Usage ### With the RKLLM API Server (VLM mode) ```bash mkdir -p ~/models/qwen3-vl-4b cd ~/models/qwen3-vl-4b git lfs install && git clone https://huggingface.co/GatekeeperZA/Qwen3-VL-4B-Instruct-RKLLM-v1.2.3 . ``` Use with [GatekeeperZA/RKLLM-API-Server](https://github.com/GatekeeperZA/RKLLM-API-Server) — the server loads both the `.rkllm` and `.rknn` files automatically when placed in the same directory. ## File Listing | File | Description | |------|-------------| | `qwen3-vl-4b-instruct_w8a8_rk3588.rkllm` | Language model weights for RK3588 NPU | | `qwen3-vl-4b-vision_rk3588.rknn` | Vision encoder for RK3588 NPU | ## Compatibility Notes - Minimum runtime: RKLLM Runtime v1.2.1 + RKNN Runtime v2.x. v1.2.3 recommended. - RKNPU driver: ≥ 0.9.6 - SoCs: RK3588 / RK3588S. Not compatible with RK3576 without reconversion. - RAM: ~5.5GB loaded. Requires 8GB+ board (16GB recommended). ## Acknowledgements - Alibaba Qwen Team for Qwen3-VL - Rockchip / airockchip for the RKLLM and RKNN toolkits - Converted by [GatekeeperZA](https://huggingface.co/GatekeeperZA)