Instructions to use minsore/pepper-1-preview with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use minsore/pepper-1-preview with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf minsore/pepper-1-preview:Q4_K_M # Run inference directly in the terminal: llama cli -hf minsore/pepper-1-preview:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf minsore/pepper-1-preview:Q4_K_M # Run inference directly in the terminal: llama cli -hf minsore/pepper-1-preview:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf minsore/pepper-1-preview:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf minsore/pepper-1-preview:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf minsore/pepper-1-preview:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf minsore/pepper-1-preview:Q4_K_M
Use Docker
docker model run hf.co/minsore/pepper-1-preview:Q4_K_M
- LM Studio
- Jan
- vLLM
How to use minsore/pepper-1-preview with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "minsore/pepper-1-preview" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "minsore/pepper-1-preview", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/minsore/pepper-1-preview:Q4_K_M
- Ollama
How to use minsore/pepper-1-preview with Ollama:
ollama run hf.co/minsore/pepper-1-preview:Q4_K_M
- Unsloth Desktop
- Pi
How to use minsore/pepper-1-preview with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf minsore/pepper-1-preview:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "minsore/pepper-1-preview:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use minsore/pepper-1-preview with Docker Model Runner:
docker model run hf.co/minsore/pepper-1-preview:Q4_K_M
- Lemonade
How to use minsore/pepper-1-preview with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull minsore/pepper-1-preview:Q4_K_M
Run and chat with the model
lemonade run user.pepper-1-preview-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use minsore/pepper-1-preview with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf minsore/pepper-1-preview:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default minsore/pepper-1-preview:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use minsore/pepper-1-preview with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf minsore/pepper-1-preview:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "minsore/pepper-1-preview:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
🌶️ Pepper 1 Preview
Lightweight instruction-following code assistant for Python, built on Qwen2.5-Coder-1.5B.
Pepper is a 1.5B parameter code assistant designed for writing functions from natural-language descriptions. It runs on consumer hardware and is one of the few publicly available 1.5B models tuned specifically for instruction-following code generation.
Part of the Minsore family · minsore.com
✨ Highlights
- 🧠 Instruction-tuned — writes functions from plain English descriptions
- 🐍 Python-focused — clean, idiomatic code without fluff
- ⚡ Real-time — designed for interactive use
- 📦 Compact — 1.5B params, ~1 GB in Q4_K_M
- 🎯 Purpose-built — trained for code generation, not chat
- 🆓 Apache 2.0 — same license as base model
📊 Benchmarks
Evaluated against 1–1.5B code models. All runs used temperature=0.0, max_tokens=512, --chat-template none.
| Benchmark | Pepper 1 Preview | Base Qwen 1.5B | Llama-3.2-1B-Code | Yi-Coder-1.5B |
|---|---|---|---|---|
| HumanEval@50 | 82.0% | 75.0% | 64.0% | 24.0% |
| MBPP@50 | 26.0% | 20.0% | 16.0% | 38.0% |
| LiveCodeBench@30 | 20.0% | 30.0% | 10.0% | 26.7% |
| BigCodeBench@30 | 23.3% | 23.3% | 13.3% | 30.0% |
📌 Key insight: Pepper 1 Preview outperforms the base model on HumanEval by +7 points — the primary benchmark for instruction-following code generation. On MBPP, both models underperform the official numbers because this evaluation uses zero-shot prompting with
--chat-template none. Yi-Coder tested without its native chat template — its HumanEval score reflects format mismatch, not model quality.
🚀 Quick Start
llama.cpp
llama-server -m pepper-1-preview.Q4_K_M.gguf \
--port 8080 \
-ngl 99 \
-c 4096 \
--chat-template chatml
Requirements: any GPU with ≥2 GB VRAM (full offload), or partial CPU offload as fallback.
Chat request
curl http://localhost:8080/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"messages": [
{"role": "user", "content": "Write a Python function that checks if a number is prime."}
],
"max_tokens": 512,
"temperature": 0.0
}'
Python
import requests
def ask(prompt, max_tokens=512):
r = requests.post("http://localhost:8080/v1/chat/completions", json={
"messages": [{"role": "user", "content": prompt}],
"max_tokens": max_tokens,
"temperature": 0.0,
})
return r.json()["choices"][0]["message"]["content"]
print(ask("Write a Python function that reverses a string."))
# → def reverse_string(s: str) -> str:
# return s[::-1]
⚙️ Recommended Settings
| Parameter | Value | Notes |
|---|---|---|
--chat-template |
chatml |
Required. Pepper expects ChatML. |
n_predict |
256–1024 | 512 is a good default |
temperature |
0.0 | Deterministic; use 0.2 for variation |
repeat_penalty |
1.1 | Prevents repetition |
-c |
4096 | Longer context available but degrades |
⚠️ Limitations
- Python-only — not trained on other languages
- No FIM support — for fill-in-the-middle tasks, use Quill
- Not an agent — does not support tool calling or multi-step planning
- Weak on algorithmic tasks — LiveCodeBench score below base Qwen
- 4K inference context — long files are truncated
- Occasional over-explanation — may include comments when only code is requested
🧬 Training Details
| Base model | Qwen2.5-Coder-1.5B-Instruct |
| Method | QLoRA fine-tuning on curated Python code data |
| Context | 4096 tokens (inference) |
📁 Files
| File | Size | Description |
|---|---|---|
pepper-1-preview.Q4_K_M.gguf |
~1 GB | Ready to use with llama.cpp |
model.safetensors |
~3 GB | Full precision (transformers) |
config.json, tokenizer.json |
— | Config for transformers |
🗺️ Roadmap
- Pepper 2 — improve MBPP and LiveCodeBench, add multi-language support
- Quill 2 — extend FIM to JS/TS/Rust
- Symphony — flagship code model (3B)
📜 License
Apache 2.0 — same as the base Qwen2.5-Coder-1.5B model.
🙏 Credits
- Base model: Qwen2.5-Coder-1.5B-Instruct by Alibaba Cloud
- Training framework: Unsloth
- Inference: llama.cpp
📬 Contact
Minsore — Ukrainian AI lab building open language models.
- 🌐 minsore.com
- 🤗 huggingface.co/Minsore
- 💬 Built by @Sollamon
⭐ If Pepper is useful, star the repo and share your results.
- Downloads last month
- 307
Model tree for minsore/pepper-1-preview
Base model
Qwen/Qwen2.5-1.5B