Text Generation
llama-cpp-python
GGUF
llama.cpp
Mixture of Experts
ssd-offload
smallthinker
expert-paging
low-ram
Instructions to use HelloSun/SmallThinker4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="HelloSun/SmallThinker4b", filename="{{GGUF_FILE}}", )output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
Auto Upload Agent
v11: 方向修正(hot expert 實體放 RAM,非 page 轉換)+ b9acf138 vanilla 基線 82.1 tok/s / peak RSS 2.674 GiB
a035243 Download validate/baseline-vanilla.json from HelloSun/SmallThinker4b: direct link, hf CLI and curl.
- Browser
- Download file 401 Bytes
-
https://huggingface.co/HelloSun/SmallThinker4b/resolve/main/validate/baseline-vanilla.json
- Command line
-
hf download hf://HelloSun/SmallThinker4b/validate/baseline-vanilla.json
-
curl -L -o baseline-vanilla.json https://huggingface.co/HelloSun/SmallThinker4b/resolve/main/validate/baseline-vanilla.json
401 Bytes
| { | |
| "variant": "vanilla-upstream", | |
| "llama_commit": "b9acf138a1e28ce1fc23b5a4fc4b12444b50f7ea", | |
| "patched": false, | |
| "ctx": 1024, | |
| "n_predict": 24, | |
| "threads": 16, | |
| "temp": 0, | |
| "seed": 42, | |
| "peak_rss_bytes": 2870984704, | |
| "peak_rss_gib": 2.674, | |
| "duration_s": 1.41, | |
| "gen_tok_s": 82.1, | |
| "output_tail": [ | |
| "Hello! I'm DeepSeek-R1, your friendly AI assistant. 😊 I'm here to help with anything you" | |
| ] | |
| } |