Text Generation
llama-cpp-python
GGUF
llama.cpp
Mixture of Experts
ssd-offload
smallthinker
expert-paging
low-ram
Instructions to use HelloSun/SmallThinker4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="HelloSun/SmallThinker4b", filename="{{GGUF_FILE}}", )output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
File size: 696 Bytes
84fbb3b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 | {
"when": "2026-10-07T16:30:49+0200",
"file": "/root/work/models/SmallThinker-4B-A0.6B-Instruct.Q4_K.gguf",
"size_mib": 512,
"fs": 4096,
"sequential": [
{
"threads": 1,
"bytes": 536870912,
"seconds": 1.518,
"mb_per_s": 337.3
},
{
"threads": 2,
"bytes": 536870912,
"seconds": 0.767,
"mb_per_s": 667.1
},
{
"threads": 4,
"bytes": 536870912,
"seconds": 0.386,
"mb_per_s": 1327.5
},
{
"threads": 8,
"bytes": 536870912,
"seconds": 0.2,
"mb_per_s": 2557.6
}
],
"random_4k": {
"count": 512,
"seconds": 0.29,
"iops": 1767.1
},
"cpu_count": 192
}
|