Text Generation
llama-cpp-python
GGUF
llama.cpp
Mixture of Experts
ssd-offload
smallthinker
expert-paging
low-ram
Instructions to use HelloSun/SmallThinker4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="HelloSun/SmallThinker4b", filename="{{GGUF_FILE}}", )output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
File size: 1,527 Bytes
625b02c f279c31 625b02c f279c31 ff122a0 84fbb3b 625b02c f279c31 625b02c f279c31 625b02c f279c31 625b02c f279c31 1e54449 f279c31 625b02c f279c31 625b02c f279c31 1e54449 f279c31 625b02c ff122a0 f279c31 ff122a0 84fbb3b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 | {
"when": "2026-10-07T16:47:29+0200",
"model": "/root/work/models/SmallThinker-4B-A0.6B-Instruct.Q4_K.gguf",
"port": 8080,
"pid": 38517,
"ram_budget_mb": 512,
"arena_mb": -1875,
"expert_slots": 252,
"chat": {
"prompt_tokens": 37,
"completion_tokens": 16,
"seconds": 3.89,
"tok_per_s": 4.1168,
"content": "A **MoE (Model Parallelism) layer** enables a neural network to",
"timings": {
"cache_n": 0,
"prompt_n": 37,
"prompt_ms": 3122.416,
"prompt_per_token_ms": 84.38962162162163,
"prompt_per_second": 11.849798361268965,
"predicted_n": 16,
"predicted_ms": 744.758,
"predicted_per_token_ms": 49.650533333333335,
"predicted_per_second": 20.140770559027224
}
},
"peak": {
"total_rss_gb": 1.0901,
"anon_rss_gb": 0.1931,
"file_rss_gb": 1.0845,
"peak_swap_gb": 0.0,
"hwm_rss_gb": 1.0901
},
"io_delta": {
"read_bytes": 1160712192,
"rchar": 0,
"write_bytes": 4096
},
"pager": {
"routing_calls": 4608,
"expert_touches": 56128,
"hits": 23134,
"misses": 2146,
"evictions": 2920,
"evicted_bytes": 6264345600,
"sweeps": 537,
"resident_bytes": 535071744,
"budget_bytes": 536870912,
"hit_rate": 0.4122,
"proc_self_read_bytes": 2439696384,
"proc_self_rchar": 11946335
},
"pager_expert_resident_mib": 510.3,
"pager_expert_budget_mib": 512.0,
"pager_expert_within_budget": true,
"pager_hit_rate": 0.4122,
"ok": true,
"model_size_mib": 2508
}
|