Instructions to use bldeaw/m1_planner_q4_k_m with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use bldeaw/m1_planner_q4_k_m with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf bldeaw/m1_planner_q4_k_m:Q4_K_M # Run inference directly in the terminal: llama cli -hf bldeaw/m1_planner_q4_k_m:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf bldeaw/m1_planner_q4_k_m:Q4_K_M # Run inference directly in the terminal: llama cli -hf bldeaw/m1_planner_q4_k_m:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf bldeaw/m1_planner_q4_k_m:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf bldeaw/m1_planner_q4_k_m:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf bldeaw/m1_planner_q4_k_m:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf bldeaw/m1_planner_q4_k_m:Q4_K_M
Use Docker
docker model run hf.co/bldeaw/m1_planner_q4_k_m:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use bldeaw/m1_planner_q4_k_m with Ollama:
ollama run hf.co/bldeaw/m1_planner_q4_k_m:Q4_K_M
- Unsloth Studio
How to use bldeaw/m1_planner_q4_k_m with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for bldeaw/m1_planner_q4_k_m to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for bldeaw/m1_planner_q4_k_m to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for bldeaw/m1_planner_q4_k_m to start chatting
- Pi
How to use bldeaw/m1_planner_q4_k_m with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bldeaw/m1_planner_q4_k_m:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "bldeaw/m1_planner_q4_k_m:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use bldeaw/m1_planner_q4_k_m with Docker Model Runner:
docker model run hf.co/bldeaw/m1_planner_q4_k_m:Q4_K_M
- Lemonade
How to use bldeaw/m1_planner_q4_k_m with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull bldeaw/m1_planner_q4_k_m:Q4_K_M
Run and chat with the model
lemonade run user.m1_planner_q4_k_m-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use bldeaw/m1_planner_q4_k_m with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bldeaw/m1_planner_q4_k_m:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default bldeaw/m1_planner_q4_k_m:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use bldeaw/m1_planner_q4_k_m with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf bldeaw/m1_planner_q4_k_m:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "bldeaw/m1_planner_q4_k_m:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf bldeaw/m1_planner_q4_k_m:Q4_K_M# Run inference directly in the terminal:
llama cli -hf bldeaw/m1_planner_q4_k_m:Q4_K_MUse pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf bldeaw/m1_planner_q4_k_m:Q4_K_M# Run inference directly in the terminal:
./llama-cli -hf bldeaw/m1_planner_q4_k_m:Q4_K_MBuild from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf bldeaw/m1_planner_q4_k_m:Q4_K_M# Run inference directly in the terminal:
./build/bin/llama-cli -hf bldeaw/m1_planner_q4_k_m:Q4_K_MUse Docker
docker model run hf.co/bldeaw/m1_planner_q4_k_m:Q4_K_MPlannerLLM
OpenAI-compatible API for bldeaw/m1_planner_q4_k_m using llama.cpp on Hugging Face Spaces.
Features
- Docker Space ready
- Persistent model storage
- Secure secret handling
- Runtime config via environment variables
- Auto model download on first start
- OpenAI-compatible endpoints
Required Hugging Face Setup
1. Create Docker Space
Create a new Hugging Face Space:
- SDK: Docker
2. Optional: Add Persistent Storage
Settings โ Storage
Recommended โ prevents re-downloading the model on every restart.
Mounted path:
/data
3. Add Variables / Secrets
Settings โ Variables and Secrets
Variables
MODEL_REPO_ID=bldeaw/m1_planner_q4_k_m
MODEL_FILE=m1_planner_q4_k_m.gguf
MODEL_REVISION=main
CTX_SIZE=2048
THREADS=2
N_GPU_LAYERS=0
MODEL_DIR=/data/models
PORT=7860
Secret (only if private repo)
HF_TOKEN=hf_xxxxx
Optional checksum
MODEL_SHA256=xxxxxxxx
Repository Files
Dockerfile
start.sh
.dockerignore
README.md
Deploy
git clone https://huggingface.co/spaces/bldeaw/planner-llm
cp Dockerfile start.sh .dockerignore README.md planner-llm/
cd planner-llm
git add .
git commit -m "production deploy"
git push origin main
API Usage
Python
from openai import OpenAI
client = OpenAI(
base_url="https://bldeaw-planner-llm.hf.space/v1",
api_key="dummy"
)
resp = client.chat.completions.create(
model="local-model",
messages=[
{"role": "user", "content": "Hello"}
]
)
print(resp.choices[0].message.content)
curl
curl https://bldeaw-planner-llm.hf.space/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "local-model",
"messages": [
{"role": "user", "content": "Hello"}
]
}'
Endpoints
| Endpoint | Method |
|---|---|
| /v1/chat/completions | POST |
| /v1/completions | POST |
| /v1/models | GET |
| / | GET |
Troubleshooting
503 Service Unavailable
Usually:
- Space building
- Model downloading (~986 MB, takes a few minutes)
- Crash during startup
Check the Logs tab.
checksum mismatch
Wrong file or corrupted download. Delete the cached file and restart.
Slow startup
Large model downloading on first run. Add persistent storage so it only downloads once.
Security
- Store tokens only in Secrets
- Use a private model repo + HF_TOKEN if needed
- Never commit
.env - Never commit
.ggufinto the Space repo
- Downloads last month
- 6
4-bit
Install (macOS, Linux)
# Start a local OpenAI-compatible server with a web UI: llama serve -hf bldeaw/m1_planner_q4_k_m:Q4_K_M# Run inference directly in the terminal: llama cli -hf bldeaw/m1_planner_q4_k_m:Q4_K_M