Text Generation
llama-cpp-python
GGUF
llama.cpp
Mixture of Experts
ssd-offload
smallthinker
expert-paging
low-ram
Instructions to use HelloSun/SmallThinker4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="HelloSun/SmallThinker4b", filename="{{GGUF_FILE}}", )output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
Auto Upload Agent
v8: 修 io_probe 的 pread 計數 bug(SSD 337/667/1328/2558 MB/s)+ launcher arena 語意修正
84fbb3b Download validate/io-probe.json from HelloSun/SmallThinker4b: direct link, hf CLI and curl.
- Browser
- Download file 696 Bytes
-
https://huggingface.co/HelloSun/SmallThinker4b/resolve/main/validate/io-probe.json
- Command line
-
hf download hf://HelloSun/SmallThinker4b/validate/io-probe.json
-
curl -L -o io-probe.json https://huggingface.co/HelloSun/SmallThinker4b/resolve/main/validate/io-probe.json
696 Bytes
| { | |
| "when": "2026-10-07T16:30:49+0200", | |
| "file": "/root/work/models/SmallThinker-4B-A0.6B-Instruct.Q4_K.gguf", | |
| "size_mib": 512, | |
| "fs": 4096, | |
| "sequential": [ | |
| { | |
| "threads": 1, | |
| "bytes": 536870912, | |
| "seconds": 1.518, | |
| "mb_per_s": 337.3 | |
| }, | |
| { | |
| "threads": 2, | |
| "bytes": 536870912, | |
| "seconds": 0.767, | |
| "mb_per_s": 667.1 | |
| }, | |
| { | |
| "threads": 4, | |
| "bytes": 536870912, | |
| "seconds": 0.386, | |
| "mb_per_s": 1327.5 | |
| }, | |
| { | |
| "threads": 8, | |
| "bytes": 536870912, | |
| "seconds": 0.2, | |
| "mb_per_s": 2557.6 | |
| } | |
| ], | |
| "random_4k": { | |
| "count": 512, | |
| "seconds": 0.29, | |
| "iops": 1767.1 | |
| }, | |
| "cpu_count": 192 | |
| } | |