v6: llama_server.sh --verify 端到端通過(預算 512MiB → expert 常駐 510MiB、swap 0) ff122a0
Auto Upload Agent commited on
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python
from llama_cpp import Llama
llm = Llama.from_pretrained(
repo_id="HelloSun/SmallThinker4b",
filename="{{GGUF_FILE}}",
)
output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)