Instructions to use litert-community/MiniCPM5-2B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- LiteRT
How to use litert-community/MiniCPM5-2B with LiteRT:
# No code snippets available yet for this library. # To use this model, check the repository files and the library's documentation. # Want to help? PRs adding snippets are welcome at: # https://github.com/huggingface/huggingface.js
- Notebooks
- Google Colab
- Kaggle
Update LlmMetadata for litert-lm 0.18.0: chat template with content parts, supports_thinking / supports_function_calling (weights unchanged)
This updates the metadata (LlmMetadata) of MiniCPM5-2B_int4.litertlm, the file I converted, in place so its URL stays the same (the size does not change). The metadata is now LiteRT-LM's official minicpm5 configuration, LlmMetadataProto.pbtext and chat_template.jinja from models/minicpm5 at 4a363c07, unchanged. It replaces the update this PR carried before, which used the template at b5e34ab1 with a content macro and set the two capability flags. Every section other than LlmMetadata is byte-identical to the published file (the unpacked section files were compared).
litert-lm 0.18.0 hands the chat template each message's content as a list of parts (models/README.md, issue 3688), so with the published template it renders the user turn empty. The file needs litert-lm 0.18.0 or newer (tested on 0.18.0): 0.17.0 passes message content as a string and this template reads parts, so the user turn renders empty.
The configuration sets the thought channel <think>\n / </think>, only the stop token ids 1 and 130073, and the model type minicpm5, so the runtime parses the model's <function name=…> tool calls. Its template answers directly unless enable_thinking is true; the published file left thinking to the model. describe prints Supports Function Call NO / Supports Thinking NO; the tool-call, thinking and tool-result checks below ran with that.
The commands that produced the file, with WORK set to an empty directory and OFF to one for the official files:
litert-lm() { uvx --from litert-lm==0.18.0 litert-lm "$@"; }
litert-lm --version
mkdir -p $OFF/minicpm5 $WORK/src $WORK/out
curl -fsSL -o $OFF/minicpm5/LlmMetadataProto.pbtext https://raw.githubusercontent.com/google-ai-edge/LiteRT-LM/4a363c0728d4461eaf05c4d87250ef3e90deb035/models/minicpm5/LlmMetadataProto.pbtext
curl -fsSL -o $OFF/minicpm5/chat_template.jinja https://raw.githubusercontent.com/google-ai-edge/LiteRT-LM/4a363c0728d4461eaf05c4d87250ef3e90deb035/models/minicpm5/chat_template.jinja
curl -fsSL -o $WORK/src/MiniCPM5-2B_int4.litertlm https://huggingface.co/litert-community/MiniCPM5-2B/resolve/8f5f487a72141710604742866bc6f67afdffc798/MiniCPM5-2B_int4.litertlm
shasum -a 256 $WORK/src/MiniCPM5-2B_int4.litertlm
litert-lm unpack $WORK/src/MiniCPM5-2B_int4.litertlm --output-dir $WORK/work
grep '^max_num_tokens' $WORK/work/LlmMetadataProto.pbtext
cp $OFF/minicpm5/LlmMetadataProto.pbtext $WORK/work/LlmMetadataProto.pbtext
grep '^max_num_tokens' $WORK/work/LlmMetadataProto.pbtext
litert-lm pack $WORK/work --chat-template $OFF/minicpm5/chat_template.jinja --output $WORK/out/MiniCPM5-2B_int4.litertlm
litert-lm unpack $WORK/out/MiniCPM5-2B_int4.litertlm --output-dir $WORK/chk --chat-template $WORK/chk.jinja && diff $WORK/chk.jinja $OFF/minicpm5/chat_template.jinja
diff -r -x LlmMetadataProto.pbtext -x model.toml $WORK/work $WORK/chk
shasum -a 256 $WORK/out/MiniCPM5-2B_int4.litertlm
Running the commands again gives a file whose sections match byte for byte while its sha256 differs: the packer writes a new uuid and creation timestamp into the container header (37 bytes in a test re-run).
Checks on litert-lm 0.18.0 (python API, CPU on a Mac, one process per check): the file answered a plain prompt, called a tool automatically and answered from its result, answered with thinking on (the reasoning on the thought channel) and off, ran two turns, read a tool result from the tool message, and answered "Paris" to a one-word system instruction (one run). On a Galaxy S26, with the prebuilt litert_lm_advanced_main from gs://litert/binaries/latest (2026-09-18 build), the file answered "Paris" to a one-word question on the CPU and on the GPU, with every text-decoder subgraph on the GPU delegate (OpenCL).
On the card, the lines on the template, the runtime requirement and the eight-question scores are rewritten (on litert-lm 0.18.0 the scores and replies match the earlier update), a section with the commands above is added, and Performance gets the Galaxy S26 rows. litertlm_manifest.json is regenerated for the new sha256 (05b78bb5…). If a new name suits the repo better, I will change the PR, and if a new export supersedes this file, I will close it.