How to use from the
Use from the
RKLLM library
# No code snippets available yet for this library.

# To use this model, check the repository files and the library's documentation.

# Want to help? PRs adding snippets are welcome at:
# https://github.com/huggingface/huggingface.js

Qwen2-0.5B-Instruct-RKLLM

RKLLM-converted Qwen2-0.5B-Instruct language-model artifacts for Rockchip RK3576 and RK3588 NPUs.

These hardware-specific .rkllm files require a compatible Rockchip RKLLM runtime. They are not Transformers checkpoints and cannot be loaded directly with Transformers, llama.cpp, or Ollama.

Base model

Conversion and variants

Toolkit version

RKLLM Toolkit: v1.3.0

Use a file built for the exact target SoC.

Target Quantization File SHA256
RK3576 W4A16 (g128) Qwen2-0.5B-Instruct_RK3576_w4a16_g128.rkllm a166cdcff5d8a2e33423d01d8bac6f7a6ad1e1130fbd0bfefdb8b69082113e87
RK3576 W8A8 Qwen2-0.5B-Instruct_RK3576_w8a8.rkllm cdccb51e23131352823af8becf9ac4acb15cb30f0f961e89d16cb0fabd003414
RK3588 W8A8 Qwen2-0.5B-Instruct_RK3588_w8a8.rkllm 24e99eeea1512d6e7f3f120f3344a510933301dfcb72bd64b24a9724ab48e1d8

The repository also includes Qwen2-0.5B-Instruct_data_quant.json, used as calibration data during conversion.

Usage

hf download HanzoHuang/Qwen2-0.5B-Instruct-RKLLM \
  RK3576/Qwen2-0.5B-Instruct_RK3576_w4a16_g128.rkllm \
  --local-dir Qwen2-0.5B-Instruct-RKLLM

Use the Qwen2 Instruct chat template with the RKLLM runtime. For Docker deployment, see Hanzo-Huang/rkllm-docker.

Limitations

These are target-specific converted artifacts. Validate quality and runtime compatibility on your Rockchip device.

Acknowledgements

Thanks to the Qwen Team, Rockchip, and the RKLLM community.

Downloads last month
36
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HanzoHuang/Qwen2-0.5B-Instruct-RKLLM

Base model

Qwen/Qwen2-0.5B
Finetuned
(594)
this model