Instructions to use tchbcb/samai-8b-M8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use tchbcb/samai-8b-M8 with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M # Run inference directly in the terminal: llama cli -hf tchbcb/samai-8b-M8:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M # Run inference directly in the terminal: llama cli -hf tchbcb/samai-8b-M8:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf tchbcb/samai-8b-M8:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf tchbcb/samai-8b-M8:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf tchbcb/samai-8b-M8:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf tchbcb/samai-8b-M8:Q4_K_M
Use Docker
docker model run hf.co/tchbcb/samai-8b-M8:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use tchbcb/samai-8b-M8 with Ollama:
ollama run hf.co/tchbcb/samai-8b-M8:Q4_K_M
- Unsloth Desktop
- Pi
How to use tchbcb/samai-8b-M8 with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "tchbcb/samai-8b-M8:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use tchbcb/samai-8b-M8 with Docker Model Runner:
docker model run hf.co/tchbcb/samai-8b-M8:Q4_K_M
- Lemonade
How to use tchbcb/samai-8b-M8 with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull tchbcb/samai-8b-M8:Q4_K_M
Run and chat with the model
lemonade run user.samai-8b-M8-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use tchbcb/samai-8b-M8 with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default tchbcb/samai-8b-M8:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use tchbcb/samai-8b-M8 with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf tchbcb/samai-8b-M8:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "tchbcb/samai-8b-M8:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
samai-8b โ M8 ็ป็ๆ้ (GGUF)
samai-8b ๆฏๅบไบ Qwen/Qwen3.5-9B๏ผbf16๏ผๆ็ปญ่ธ้ฆๅขๅผบ็ๅฏน่ฏ/ๆง่กๆจกๅใๆฌไปๅบไบคไป M8ใagent ๆง่ก่ฝๅๅขๅผบใ่ฝฎ็ๆ็ป q4_k_m ้ๅๆ้๏ผไธบ็ณปๅๅทฅไฝ็็ป็ไบคไปใ
่ฏๆต็ปๆ
| ่ฏๆต้ | M7 ๅบ็บฟ | M8 ็ป็ | ๆๅ |
|---|---|---|---|
| agent60 A ๅ่ทณ | 100.0 | 100.0 | โ |
| agent60 B ๅคๅทฅๅ ท | 73.3 | 100.0 | +26.7 |
| agent60 C ๅค่ฝฎ้พๅผ | 100.0 | 93.3 | -6.7 |
| agent60 D ่ดๅ็ด็ญ | 80.0 | 93.3 | +13.3 |
| agent60 ๆปๅ | 86.7 | 96.7 | +10.0 |
| knight15 ่ฝๅๅๅฝ | 15/15 | 15/15 | ้ถๅ้ |
ๅฎ้ถๅฝไธญ๏ผๅคๅทฅๅ ท็ผๆ๏ผB +26.7๏ผไธ่ดๅๆ็ญ็ด็ญ๏ผD +13.3๏ผๅๅๆพ่ๆๅ๏ผknight15 ๆปกๅ็กฎ่ฎคๆ ธๅฟ่ฝๅ้ถๅ้ใ
ๆไปถ
| ๆไปถ | ๅคงๅฐ | sha256 |
|---|---|---|
samai8b_m8_q4_k_m.gguf |
5.78 GB (5,780,090,720 B) | 9566d8a7d85782b2284ee970d82306dfd33fb52df4b35e42a722a3d63c94f818 |
ๆบฏๆบไธ่ฎญ็ป้ ๆน
- ๅบๅบง๏ผ
Qwen/Qwen3.5-9Bbf16๏ผ้ๅฏนๆ ก้ชๆตๅผๅๅนถ LoRA๏ผ128/128 lora pairs ๅ จๅฝไธญ๏ผ - ่ฎญ็ป๏ผQLoRA๏ผnf4, r=32, ฮฑ=64๏ผ๏ผๆๅธๅณ็จๅบ่ๅผๅๆ็ๅ็ฐ agent ๆไปคๆฐๆฎ๏ผๅ่ทณ/ๅคๅทฅๅ ท/ๅค่ฝฎ้พๅผ/่ดๅ็ด็ญ๏ผๅ ฑ 3200 ๆก๏ผๅซๅทฅๅ ทๅฏ่พพๆง็บข็บฟๆ ก้ชไธ็็ญๆก่ดๆ ทๆฌ๏ผ
- ้ๅ้พ๏ผๅๅนถ bf16 โ
convert_hf_to_gguf.pyf16 โllama-quantizeq4_k_m - ่ฏๆต๏ผ60 ้ข agent ่ฏๆต๏ผT4 ็ๆบ๏ผ+ knight15 ่ฝๅๅๅฝ้จ๏ผๆฅๅ่ง
m8_final_report.json/m8_knight15_report.json/m8_baseline_report.json
็จๆณ (llama.cpp)
llama-server -m samai8b_m8_q4_k_m.gguf -ngl 99 -c 2048 --port 8080
q4_k_m ๅจ T4 (SM75) ไธ้ช่ฏ้่ฟ๏ผCPU ๆจกๅผ็ๆจ็ๅ็ไบฆ้่ฟใ
PK: samai-8b M8 vs Ternary-Bonsai-2-27B (prism-ml)
ๅ็ไฝๅฏนๅณ (q4_k_m 5.78G vs PTQ1_0 5.95G), ๅ จๆฐ 48 ้ข agent ไปปๅก้ถ่ก (้ถ่ฎญ็ปๆณๆผ, ๅ็ฐ: ๅ่ทณ/ๅคๅทฅๅ ท/ๅค่ฝฎ้พๅผ/่ดๅ็ด็ญ), ๅ harness / ๅๆ server ้ ็ฝฎ / greedy ๅคๅ:
| ็ฐ | samai-8b M8 | Bonsai-2-27B | delta |
|---|---|---|---|
| A ๅ่ทณ | 58.3% | 58.3% | 0 |
| B ๅคๅทฅๅ ท | 50.0% | 41.7% | +8.3 |
| C ๅค่ฝฎ้พๅผ | 41.7% | 41.7% | 0 |
| D ่ดๅ็ด็ญ | 100.0% | 100.0% | 0 |
| ๆปๅ | 62.5% | 60.4% | +2.1 |
่ฃๅณ: M8 ่ๅบ โ ไปฅ 9B/5.78G ๆๅนณๅนถๅจๆปๅไธ้้ข head-to-head (1 ่ 0 ่ด) ไธๅ่ฟ 27B/5.95G ็ ternary ๅ็บงๅฏนๆ, ๅคๅทฅๅ
ท็ผๆๆฏๅณ่็ฐใ้้ขๆ็ป่ง pk_summary.json, ้ขๅบ่ง m8_pk_eval.jsonlใ
- Downloads last month
- -
4-bit