Msma12's picture
|
download
raw
4.17 kB
---
license: apache-2.0
language:
- en
- zh
base_model:
- openbmb/MiniCPM5-1B
pipeline_tag: text-generation
library_name: litert
tags:
- minicpm
- minicpm5
- litert
- tflite
- on-device
- edge-ai
---
# MiniCPM5-1B (LiteRT-LM)
This repository hosts the [**LiteRT-LM**](https://ai.google.dev/edge/litert-lm) (LiteRT formerly known as TensorFlow Lite) version of **MiniCPM5-1B**, optimized for fully on-device inference on mobile and edge hardware.
---
## Available Models
* **`minicpm_dynamic_wi8_afp32_gpu_opt.litertlm`**: This model features dynamic weight-only INT8 quantization (wi8) with FP32 activations (afp32), heavily optimized for GPU execution.
## What is MiniCPM?
**MiniCPM5-1B** is the first model in the **MiniCPM5** series from [OpenBMB](https://huggingface.co/openbmb). It is a dense **1B-parameter** Transformer built specifically for **on-device, local, and resource-constrained deployment**, while reaching **1B-class open-source SOTA** in its size class.
### Highlights
- ๐Ÿ† **1B-class open-source SOTA** โ€” strongest in tool use, code generation, and difficult reasoning among comparable open models.
- ๐Ÿง  **Hybrid Reasoning** โ€” a single checkpoint serves as both a fast assistant and a deliberate reasoner via a built-in `<think>` template (`enable_thinking`).
- ๐Ÿ“ **Long context** โ€” native **131,072**-token context length.
- ๐Ÿ“ฑ **Built for the edge** โ€” compact footprint designed for local assistants, coding agents, and tool-use workflows.
### Model Information
| Item | Value |
| --- | --- |
| Type | Causal Language Model |
| Architecture | Standard `LlamaForCausalLM` |
| Parameters | 1,080,632,832 (~1B) |
| Non-Embedding Parameters | 679,552,512 |
| Layers | 24 |
| Attention Heads (GQA) | 16 (Q) / 2 (KV) |
| Context Length | 131,072 |
---
## Use the model
### Edge Gallery App (Android)
1. **Get the App**: Install the [app](https://play.google.com/store/apps/details?id=com.google.ai.edge.gallery&pli=1) from Google Play or download the latest APK from the [GitHub releases page](https://github.com/google-ai-edge/gallery/releases).
2. **Importing the Model**: Navigate to the **Model manager** within the app and click the **"+" (plus)** icon in the bottom-right corner. Two options will appear:
* **Import from HF (Recommended)**: Select this option, and a dialog box will appear showing an example Hugging Face model URL. Enter the HF link for the desired `.litertlm` model and click submit. The model will then appear in your list, and you can proceed to download it (a Hugging Face account login is required).
* **From local model file**: First, download the `.litertlm` model directly to your Android device, OR download it to your computer and push it via ADB (e.g., `adb push minicpm_dynamic_wi8_afp32_gpu_opt.litertlm /sdcard/Download/`). Then, select this option, choose the downloaded file from your storage, configure your preferred parameters, and tap **"Import"**.
For full details on importing models and other features, see the [Edge Gallery App Wiki](https://github.com/google-ai-edge/gallery/wiki).
To build the demo app from source, please follow the [instructions](https://github.com/google-ai-edge/gallery/blob/main/README.md) from the GitHub repository.
### Try It (Desktop/CLI)
Install `uv` and run the model directly from the LiteRT-LM command line:
```bash
uv tool install litert-lm
uvx litert-lm run --from-huggingface-repo=litert-community/MiniCPM5-1B minicpm_dynamic_wi8_afp32_gpu_opt.litertlm --prompt="What is the capital of France?"
```
## Links
- ๐Ÿค— Original model (BF16): [openbmb/MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B)
- ๐Ÿ“ฆ GitHub: [OpenBMB/MiniCPM](https://github.com/OpenBMB/MiniCPM)
- ๐Ÿ› ๏ธ LiteRT docs: [ai.google.dev/edge/litert](https://ai.google.dev/edge/litert)
---
## License
Released under the **Apache-2.0 License**, consistent with the upstream [openbmb/MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B).
## Citation
```bibtex
@article{minicpm4,
title={MiniCPM4: Ultra-efficient LLMs on end devices},
author={MiniCPM, Team},
journal={arXiv preprint arXiv:2506.07900},
year={2025}
}
```

Xet Storage Details

Size:
4.17 kB
ยท
Xet hash:
34e4c47dc7282b4d31dcffad0454d8cc6607e85f2579a371732a04fe6a457aa6

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.