v10: model card 加上完整實測數據(記憶體/速度/SSD/權重組成);修 pager expert 總量高估 7% 的 bug f279c31
Auto Upload Agent commited on
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python
from llama_cpp import Llama
llm = Llama.from_pretrained(
repo_id="HelloSun/SmallThinker4b",
filename="{{GGUF_FILE}}",
)
output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)