majentik's picture
Card accuracy sweep: honest brand labeling, remove dead links, upstream KV tip
b046ab5 verified
|
Raw
History Blame Contribute Delete
1.62 kB
---
base_model: openbmb/MiniCPM5-1B-Base
license: apache-2.0
library_name: gguf
pipeline_tag: text-generation
tags:
- gguf
- quantized
- rotorquant
- llama.cpp
---
> [!TIP]
> **KV-cache quantization (upstream, no fork needed):** llama.cpp/Ollama cover
> this natively — `-ctk q8_0 -ctv q8_0` (~half KV memory, negligible quality
> loss) or `-ctk q4_0 -ctv q4_0` (~quarter memory, small quality cost). In
> Ollama: `OLLAMA_KV_CACHE_TYPE=q8_0` with `OLLAMA_FLASH_ATTENTION=1`.
# MiniCPM5-1B-Base — RotorQuant GGUF Q8_0
[`openbmb/MiniCPM5-1B-Base`](https://huggingface.co/openbmb/MiniCPM5-1B-Base) quantized pack, published as `MiniCPM5-1B-Base-RotorQuant-GGUF-Q8_0`.
## Method
llama.cpp Q8_0 quantization.
## Release line
Released under the **RotorQuant** line. RotorQuant and TurboQuant are this project's release labels for this pack, not distinct quantization algorithms — both brand repos carry byte-identical weights. No brand-specific speedup is claimed or measured.
## Modality
`pipeline_tag: text-generation`.
This is a llama.cpp GGUF conversion of the text tower; no modality beyond `pipeline_tag` above is claimed or included.
## License
This pack is a **derivative** of [`openbmb/MiniCPM5-1B-Base`](https://huggingface.co/openbmb/MiniCPM5-1B-Base); all credit for the original model, training, and weights belongs to the upstream authors. This repo republishes a quantized conversion of those weights only.
Governed by the [apache-2.0](https://www.apache.org/licenses/LICENSE-2.0). See the upstream repo and the linked license for the full terms — no license text is reproduced here.