Text Generation
llama-cpp-python
GGUF
llama.cpp
Mixture of Experts
ssd-offload
smallthinker
expert-paging
low-ram
Instructions to use HelloSun/SmallThinker4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="HelloSun/SmallThinker4b", filename="{{GGUF_FILE}}", )output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
File size: 401 Bytes
a035243 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 | {
"variant": "vanilla-upstream",
"llama_commit": "b9acf138a1e28ce1fc23b5a4fc4b12444b50f7ea",
"patched": false,
"ctx": 1024,
"n_predict": 24,
"threads": 16,
"temp": 0,
"seed": 42,
"peak_rss_bytes": 2870984704,
"peak_rss_gib": 2.674,
"duration_s": 1.41,
"gen_tok_s": 82.1,
"output_tail": [
"Hello! I'm DeepSeek-R1, your friendly AI assistant. 😊 I'm here to help with anything you"
]
} |