Spaces:
Sleeping
Sleeping
metadata
title: PlannerLLM
emoji: 🧠
colorFrom: blue
colorTo: green
sdk: docker
pinned: false
short_description: OpenAI-compatible planner model API via llama.cpp
PlannerLLM
OpenAI-compatible API for bldeaw/m1_planner_q4_k_m using llama.cpp on Hugging Face Spaces.
Features
- Docker Space ready
- Persistent model storage
- Secure secret handling
- Runtime config via environment variables
- Auto model download on first start
- OpenAI-compatible endpoints
Required Hugging Face Setup
1. Create Docker Space
Create a new Hugging Face Space:
- SDK: Docker
2. Optional: Add Persistent Storage
Settings → Storage
Recommended — prevents re-downloading the model on every restart.
Mounted path:
/data
3. Add Variables / Secrets
Settings → Variables and Secrets
Variables
MODEL_REPO_ID=bldeaw/m1_planner_q4_k_m
MODEL_FILE=m1_planner_q4_k_m.gguf
MODEL_REVISION=main
CTX_SIZE=2048
THREADS=2
N_GPU_LAYERS=0
MODEL_DIR=/data/models
PORT=7860
Secret (only if private repo)
HF_TOKEN=hf_xxxxx
Optional checksum
MODEL_SHA256=xxxxxxxx
Repository Files
Dockerfile
start.sh
.dockerignore
README.md
Deploy
git clone https://huggingface.co/spaces/bldeaw/planner-llm
cp Dockerfile start.sh .dockerignore README.md planner-llm/
cd planner-llm
git add .
git commit -m "production deploy"
git push origin main
API Usage
Python
from openai import OpenAI
client = OpenAI(
base_url="https://bldeaw-planner-llm.hf.space/v1",
api_key="dummy"
)
resp = client.chat.completions.create(
model="local-model",
messages=[
{"role": "user", "content": "Hello"}
]
)
print(resp.choices[0].message.content)
curl
curl https://bldeaw-planner-llm.hf.space/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "local-model",
"messages": [
{"role": "user", "content": "Hello"}
]
}'
Endpoints
| Endpoint | Method |
|---|---|
| /v1/chat/completions | POST |
| /v1/completions | POST |
| /v1/models | GET |
| / | GET |
Troubleshooting
503 Service Unavailable
Usually:
- Space building
- Model downloading (~986 MB, takes a few minutes)
- Crash during startup
Check the Logs tab.
checksum mismatch
Wrong file or corrupted download. Delete the cached file and restart.
Slow startup
Large model downloading on first run. Add persistent storage so it only downloads once.
Security
- Store tokens only in Secrets
- Use a private model repo + HF_TOKEN if needed
- Never commit
.env - Never commit
.ggufinto the Space repo