MiniCPM3-4B-RKLLM

RKLLM-converted MiniCPM3-4B language-model artifacts for Rockchip RK3576 and RK3588 NPUs.

These are hardware-specific .rkllm binaries, not Transformers checkpoints. They require a compatible Rockchip RKLLM runtime and cannot be loaded directly with Transformers, llama.cpp, or Ollama.

Base model

Conversion and variants

Toolkit version

RKLLM Toolkit: v1.3.0

Use a file built for the exact target SoC.

Target Quantization File SHA256
RK3576 W4A16 (g128) MiniCPM3-4B_RK3576_w4a16_g128.rkllm ec05211d17c4bb05a685588ff335608233b59722f6ecc296976d61dbc88c4128
RK3576 W8A8 MiniCPM3-4B_RK3576_w8a8.rkllm 3795c6a22e5f12de0c4243fc187da55f516b530de3125734874d340856160aec
RK3588 W8A8 MiniCPM3-4B_RK3588_w8a8.rkllm 3815a59593df7aebda90f86844caf2fd625d9d0e996e35aa31fd74790d10b130

The repository also includes MiniCPM3-4B_data_quant.json, used as calibration data during conversion.

Usage

hf download HanzoHuang/MiniCPM3-4B-RKLLM \
  RK3576/MiniCPM3-4B_RK3576_w4a16_g128.rkllm \
  --local-dir MiniCPM3-4B-RKLLM

Run the downloaded file with the RKLLM runtime and the upstream MiniCPM3 chat template. For Docker deployment, see Hanzo-Huang/rkllm-docker.

Limitations

These artifacts are target-specific conversions. Validate quality, memory use, and runtime compatibility on your own hardware.

Acknowledgements

Thanks to OpenBMB, Rockchip, and the RKLLM community.

Downloads last month
9
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HanzoHuang/MiniCPM3-4B-RKLLM

Finetuned
(2)
this model