DIRKIEdev/think / README.md
DIRKIEdev's picture
|
download
raw
4.21 kB
---
license: apache-2.0
base_model: GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking
base_model_relation: quantized
language:
- en
- zh
pipeline_tag: text-generation
library_name: gguf
tags:
- gguf
- llama.cpp
- quantized
- minicpm5
- thinking
- fable5
- coding
- instruction-following
---
<p align="center">
<img src="assets/banner.png" alt="MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF" width="100%"/>
</p>
# MiniCPM5-1B-Claude-Opus-Fable5-Thinking-GGUF
> **πŸ“’ V2.0 is available** β€” We have released an updated model with **enhanced tool-calling** capabilities. Welcome to try the new version:
> - Transformers: [MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking](https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking)
> - GGUF: [MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF](https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-V2-Thinking-GGUF)
GGUF quantizations of **[MiniCPM5-1B-Claude-Opus-Fable5-Thinking](https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking)** for [llama.cpp](https://github.com/ggml-org/llama.cpp), Ollama, LM Studio, jan, KoboldCpp, and other GGUF runtimes.
[δΈ­ζ–‡θ―΄ζ˜Ž](./README-cn.md)
This repository provides local-deployment builds of a 1B **Thinking** model fine-tuned on **Fable 5** data atop [openbmb/MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B). The GGUF files embed MiniCPM5's native chat template for llama.cpp-compatible runtimes.
Transformers checkpoint: **[MiniCPM5-1B-Claude-Opus-Fable5-Thinking](https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking)**
---
## Files
| File | Quant | Size | Notes |
|---|---|---|---|
| `MiniCPM5-1B-Claude-Opus-Fable5-Thinking-Q4_K_M.gguf` | Q4_K_M | ~657 MB | smallest footprint |
| `MiniCPM5-1B-Claude-Opus-Fable5-Thinking-Q5_K_M.gguf` | Q5_K_M | ~751 MB | balanced quality / size |
| `MiniCPM5-1B-Claude-Opus-Fable5-Thinking-Q8_0.gguf` | Q8_0 | ~1.1 GB | **recommended default** |
| `MiniCPM5-1B-Claude-Opus-Fable5-Thinking-F16.gguf` | F16 | ~2.1 GB | full-precision conversion base |
**Q8_0** is the recommended default quant for this 1B model.
---
## Quick start
### llama.cpp (`llama-cli`)
```bash
llama-cli \
-m MiniCPM5-1B-Claude-Opus-Fable5-Thinking-Q8_0.gguf \
-p "Write a Python function to merge two sorted lists." \
-n 512 \
--temp 0.9 --top-p 0.95 \
-c 8192
```
> The model supports up to **128K tokens** (131,072) per `config.json`. Set `-c` according to your available VRAM/RAM.
### llama.cpp server
```bash
llama-server \
-m MiniCPM5-1B-Claude-Opus-Fable5-Thinking-Q8_0.gguf \
-c 8192 --port 8080
```
### LM Studio / jan / KoboldCpp
Load any `.gguf` file from this repository. The MiniCPM5 chat template is embedded in the GGUF metadata.
---
## Sampling recommendations
Generation defaults are inherited from **[MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B)**:
| Mode | Params |
|---|---|
| **Think** (default) | `temperature=0.9, top_p=0.95` |
| **No Think** | `temperature=0.7, top_p=0.95`, `enable_thinking=False` |
---
## Capabilities
- **Fable 5 fine-tune** β€” post-trained on Fable 5 data
- **Coding** β€” code generation, debugging, and software-engineering workflows
- **Instruction following** β€” more reliable adherence to user prompts and task constraints
- **Thinking mode** β€” chain-of-thought reasoning; MiniCPM5 chat template baked into the GGUF
- **Long context** β€” up to **128K tokens** (131,072 tokens per upstream `config.json`)
---
## Limitations
- **Thinking outputs** β€” the model may emit reasoning blocks before the final answer
- **1B scale** β€” lightweight local deployment; not frontier-scale
- **Runtime context** β€” actual usable context depends on your GGUF runtime and hardware limits
---
## Provenance & licensing
**Apache-2.0**, inherited from [MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B).
## Acknowledgements
- Base model: [OpenBMB / MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B)
- Transformers checkpoint: [MiniCPM5-1B-Claude-Opus-Fable5-Thinking](https://huggingface.co/GnLOLot/MiniCPM5-1B-Claude-Opus-Fable5-Thinking)
- Quantization: [llama.cpp](https://github.com/ggml-org/llama.cpp)

Xet Storage Details

Size:
4.21 kB
Β·
Xet hash:
8b62b7ccf45e38694f6bb544404f1d68663457db9e46f01279b0e15545857bb3

Xet efficiently stores files, intelligently splitting them into unique chunks and accelerating uploads and downloads. More info.