Text Generation
llama-cpp-python
GGUF
llama.cpp
Mixture of Experts
ssd-offload
smallthinker
expert-paging
low-ram
Instructions to use HelloSun/SmallThinker4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- llama-cpp-python
How to use HelloSun/SmallThinker4b with llama-cpp-python:
# !pip install llama-cpp-python from llama_cpp import Llama llm = Llama.from_pretrained( repo_id="HelloSun/SmallThinker4b", filename="{{GGUF_FILE}}", )output = llm( "Once upon a time,", max_tokens=512, echo=True ) print(output)
- Notebooks
- Google Colab
- Kaggle
File size: 1,264 Bytes
faea13d | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 | #!/usr/bin/env python3
"""丟掉模型檔在系統 page cache 裡的頁(POSIX_FADV_DONTNEED)。
為什麼需要:量 SSD 讀取量之前一定要做,否則「上一輪留在核心 page cache」
會讓數字虛高(看起來像是 SSD 變快了,其實是記憶體)。
用法:
python3 tools/drop_model_cache.py <model.gguf> [more files...]
"""
import os
import sys
POSIX_FADV_DONTNEED = 4
SYNC = 2
def drop(path: str) -> tuple[bool, str]:
try:
fd = os.open(path, os.O_RDONLY)
except OSError as e:
return False, str(e)
try:
size = os.fstat(fd).st_size
os.posix_fadvise(fd, 0, size, POSIX_FADV_DONTNEED)
os.fsync(fd)
return True, f"{size} bytes"
except (AttributeError, OSError) as e:
return False, str(e)
finally:
os.close(fd)
def main() -> int:
if len(sys.argv) < 2:
print(__doc__)
return 2
rc = 0
for p in sys.argv[1:]:
if not os.path.exists(p):
print(f"[drop] 找不到 {p}")
rc = 1
continue
ok, info = drop(p)
print(f"[drop] {'ok ' if ok else 'fail'} {p} {info}")
if not ok:
rc = 1
return rc
if __name__ == "__main__":
sys.exit(main()) |