Instructions to use HanzoHuang/Llama-3.2-1B-Instruct-RKLLM with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- RKLLM
How to use HanzoHuang/Llama-3.2-1B-Instruct-RKLLM with RKLLM:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
File size: 2,237 Bytes
c1264ae d5c5e40 c1264ae d5c5e40 094237b d5c5e40 094237b d5c5e40 094237b d5c5e40 094237b d5c5e40 094237b d5c5e40 094237b d5c5e40 67adf08 d5c5e40 094237b d5c5e40 094237b d5c5e40 094237b d5c5e40 094237b d5c5e40 094237b d5c5e40 094237b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 | ---
license: llama3.2
base_model: meta-llama/Llama-3.2-1B-Instruct
pipeline_tag: text-generation
library_name: rkllm
tags:
- rkllm
- rockchip
- rk3576
- rk3588
- llama
- llama-3.2
---
# Llama-3.2-1B-Instruct-RKLLM
RKLLM-converted Llama 3.2 1B Instruct language-model artifacts for Rockchip RK3576 and RK3588 NPUs.
These hardware-specific `.rkllm` files require a compatible Rockchip RKLLM runtime. They are not Transformers checkpoints and cannot be loaded directly with Transformers, llama.cpp, or Ollama.
## Base model
- Upstream model: [meta-llama/Llama-3.2-1B-Instruct](https://huggingface.co/meta-llama/Llama-3.2-1B-Instruct)
- License: Llama 3.2 Community License
- Model type: LLM (text only)
The upstream model is gated. Review Meta's license and acceptable-use policy before use or redistribution.
## Conversion and variants
### Toolkit version
**RKLLM Toolkit: v1.3.0**
Use a file built for the exact target SoC.
| Target | Quantization | File | SHA256 |
| --- | --- | --- | --- |
| RK3576 | W4A16 (g128) | [Llama-3.2-1B-Instruct_RK3576_w4a16_g128.rkllm](RK3576/Llama-3.2-1B-Instruct_RK3576_w4a16_g128.rkllm) | `bd6cbbf84e90b63d451e672a542858a5a6aa15ad2b45f83e1b5dbcd944ac3018` |
| RK3576 | W8A8 | [Llama-3.2-1B-Instruct_RK3576_w8a8.rkllm](RK3576/Llama-3.2-1B-Instruct_RK3576_w8a8.rkllm) | `9c4d103f7c486740ba75605c61a51b80d170df067645360ffad8464b50c764df` |
| RK3588 | W8A8 | [Llama-3.2-1B-Instruct_RK3588_w8a8.rkllm](RK3588/Llama-3.2-1B-Instruct_RK3588_w8a8.rkllm) | `4403aec1c018c7ebf29c95730336e848aff8824e3be3bd6be7db44b4a4973bad` |
The repository also includes `Llama-3.2-1B-Instruct_data_quant.json`, used as calibration data during conversion.
## Usage
```bash
hf download HanzoHuang/Llama-3.2-1B-Instruct-RKLLM \
RK3576/Llama-3.2-1B-Instruct_RK3576_w4a16_g128.rkllm \
--local-dir Llama-3.2-1B-Instruct-RKLLM
```
Use the Llama 3.2 Instruct chat template with the RKLLM runtime. For Docker deployment, see [Hanzo-Huang/rkllm-docker](https://github.com/Hanzo-Huang/rkllm-docker).
## Limitations
These are target-specific converted artifacts. Validate quality, memory use, and runtime compatibility on your own Rockchip device.
## Acknowledgements
Thanks to Meta, Rockchip, and the RKLLM community.
|