| --- |
| title: Inference |
| emoji: ⚡ |
| colorFrom: gray |
| colorTo: indigo |
| sdk: docker |
| pinned: false |
| --- |
| |
| # CPU Inference Backend |
|
|
| This Space is a backend-only OpenAI-compatible inference server built on `llama.cpp` and optimized for CPU-only Hugging Face Docker Spaces. |
|
|
| It automatically downloads the target Hugging Face model from `MODEL_NAME`, stores it under `/data/models/model.gguf`, and starts `llama-server` with OpenAI-compatible endpoints. |
|
|
| ## Environment variables |
|
|
| Set these in the Space settings: |
|
|
| - `MODEL_NAME`: Hugging Face repo id that contains one or more GGUF files |
| - `MODEL_FILE`: optional exact GGUF filename inside `MODEL_NAME`; useful for testing faster/lower quality quants |
| - `HF_TOKEN`: Hugging Face token used for the download |
| - `API_PASSWORD`: bearer token required by the API |
|
|
| Optional tuning variables: |
|
|
| - `CTX_SIZE`: context size, default `4096` (reduce to 2048-4096 for small models to improve speed) |
| - `THREADS`: CPU thread count, default auto-detect via cgroup CPU quota, then `nproc` |
| - `OMP_PROC_BIND`, `OMP_PLACES`, `OMP_WAIT_POLICY`: OpenMP CPU placement controls, default `FALSE`, `cores`, and `PASSIVE` |
| - `THREADS_BATCH`: batch thread count, default matches THREADS |
| - `BATCH_SIZE`: prompt batch size, default 128 for balanced mode |
| - `UBATCH_SIZE`: micro-batch size, default 128 for balanced mode |
| - `CACHE_TYPE_K`: KV cache type for keys, default `f16` (use `q4_0` for memory savings) |
| - `CACHE_TYPE_V`: KV cache type for values, default `f16` (use `q4_0` for memory savings) |
| - `PERF_PROFILE`: `low_latency`, `balanced` (default), or `throughput` |
| - `REASONING`: reasoning mode, accepts `True`, `False`, or `auto` and maps to llama.cpp `--reasoning` |
| - `LANGSEARCH_API_KEY`: optional API key for LangSearch web search tool (free tier: 1 req/sec, 60/min, 1000/day) |
| - `ENABLE_TOOLS`: set to `1` to enable llama.cpp shell tool support when `LANGSEARCH_API_KEY` is present, default `0` |
| - `HTTP_THREADS`: HTTP server worker threads, default `1` for personal single-request inference |
| - `LOG_VERBOSITY`: llama.cpp log verbosity, default `1` to reduce runtime logging overhead |
| - `MMAP`: set to `1` to use memory-mapped model loading, default `0` |
| - `FLASH_ATTN`: set to `0` to disable flash attention if a model hangs or fails during startup, default `1` |
| - `NO_WARMUP`: set to `1` to skip model warmup and save startup time |
| - `PORT`: listen port, default `7860` |
|
|
| ## Endpoints |
|
|
| The server exposes: |
|
|
| - `/v1/chat/completions` |
| - `/v1/completions` |
| - `/v1/models` |
|
|
| Streaming is supported by `llama-server`. Bearer token authentication uses `--api-key` directly. |
|
|
| ## Build and runtime |
|
|
| The Docker image uses a multi-stage build on Debian bookworm. The builder stage installs `git`, `build-essential`, `cmake`, and OpenBLAS, clones `llama.cpp`, and compiles `llama-server` with `-Ofast`, `-march=native`, `-flto`, OpenBLAS, and native CPU optimizations. The runtime stage keeps only Python, OpenBLAS/OpenMP runtime libraries, the compiled server, and the downloader. |
|
|
| Model downloads and HF cache live on the `/data` bucket so restarts do not redownload the model. |
|
|
| ## Performance Optimization |
|
|
| For small 2B models on CPU-only inference (~2-4 t/s baseline), optimize these settings: |
|
|
| **High Impact:** |
| - The server forces `--parallel 1`, so only one request runs at a time and all CPU is focused on the active query |
| - Runtime uses `--no-mmap`, `--flash-attn on`, `--threads-http 1`, and low log verbosity to favor single-request inference over serving overhead |
| - OpenBLAS is enabled because benchmark results on the Space showed BLAS-off hurt prompt processing and did not recover generation speed |
| - `BATCH_SIZE` & `UBATCH_SIZE`: Use **64-256** for faster prompt ingestion; tiny values like 8 cause many small prompt batches |
| - `PERF_PROFILE`: Use **`balanced`** for normal personal use; `throughput` only changes prompt batch sizing now |
|
|
| **Medium Impact:** |
| - `NO_WARMUP=1`: Skip model warmup to reduce startup latency by ~0.2s |
| - `CACHE_TYPE_K`, `CACHE_TYPE_V`: Try `q4_0` instead of `f16` to reduce KV cache bandwidth and memory use at longer contexts |
| - `ENABLE_TOOLS`: Leave as `0` unless you need tool calling. Built-in shell tools add prompt/tool overhead and should not be exposed publicly. |
| - `MODEL_FILE`: If the repo has multiple GGUFs, test a lighter quant such as `Q4_0`, `Q4_K_S`, or `Q3_K_M` for speed/quality tradeoffs. |
| - Compatibility: if a GGUF hangs during startup, try `FLASH_ATTN=0`, `CACHE_TYPE_K=f16`, `CACHE_TYPE_V=f16`, and then `MMAP=1`. |
|
|
| **Example for 2B model (CPU-only):** |
| ``` |
| CTX_SIZE=2048 |
| BATCH_SIZE=128 |
| UBATCH_SIZE=128 |
| PERF_PROFILE=balanced |
| REASONING=False |
| ENABLE_TOOLS=0 |
| CACHE_TYPE_K=q4_0 |
| CACHE_TYPE_V=q4_0 |
| ``` |
|
|
| Expected: faster first-token latency from prompt ingestion. Generation speed is still CPU-bound on the free tier, so a 2B Q4 model will usually remain in the low single-digit tokens/sec range. |
|
|
| ## Web Search Tool (Standalone) |
|
|
| When `ENABLE_TOOLS=1` and `LANGSEARCH_API_KEY` is set, `start.sh` enables llama.cpp's shell tool so the model can launch the local LangSearch wrapper. |
|
|
| llama.cpp does not expose arbitrary external HTTP APIs as first-class built-in tools, so LangSearch is still implemented as a local wrapper that runs through the shell tool. |
|
|
| **Use cases:** |
| - Call from your application layer to augment model responses |
| - Pre-process queries before sending to model |
| - Post-process to fetch real-time data for specific topics |
|
|
| **API Usage (from application):** |
|
|
| ```python |
| import requests |
| import json |
| |
| response = requests.post( |
| "http://localhost:7860/v1/chat/completions", |
| headers={"Authorization": "Bearer YOUR_API_KEY"}, |
| json={ |
| "model": "model", |
| "messages": [{"role": "user", "content": "What's the latest news?"}], |
| "temperature": 0.7 |
| } |
| ) |
| |
| latest_response = response.json()['choices'][0]['message']['content'] |
| ``` |
|
|
| Then call the search tool from your backend or from the shell tool: |
|
|
| ```bash |
| python3 /app/search_tool.py "latest AI news" 5 true noLimit |
| ``` |
|
|
| Combine results before returning to user. |
|
|
| **Command-line (Docker container):** |
|
|
| Examples: |
| ```bash |
| python3 /app/search_tool.py "latest AI news" 5 true noLimit |
| python3 /app/search_tool.py "Python 3.13 release" 3 false oneWeek |
| ``` |
|
|
| Parameters: |
| - `query`: search string (required) |
| - `count`: max results 1-10 (default: 5) |
| - `summary`: include summaries (default: true) |
| - `freshness`: `oneDay`, `oneWeek`, `oneMonth`, `oneYear`, `noLimit` (default: `noLimit`) |
|
|
| Output is JSON with search results, URLs, snippets, and optional summaries. |
|
|
| ## curl examples |
|
|
| ```bash |
| curl http://localhost:7860/v1/models \ |
| -H "Authorization: Bearer $API_PASSWORD" |
| ``` |
|
|
| ```bash |
| curl http://localhost:7860/v1/chat/completions \ |
| -H "Authorization: Bearer $API_PASSWORD" \ |
| -H "Content-Type: application/json" \ |
| -d '{ |
| "model": "model", |
| "messages": [ |
| {"role": "user", "content": "Write a one-sentence summary of llama.cpp."} |
| ], |
| "stream": false |
| }' |
| ``` |
|
|
| ```bash |
| curl http://localhost:7860/v1/completions \ |
| -H "Authorization: Bearer $API_PASSWORD" \ |
| -H "Content-Type: application/json" \ |
| -d '{ |
| "model": "model", |
| "prompt": "Explain KV cache in one paragraph.", |
| "stream": false |
| }' |
| ``` |
|
|
| ## OpenAI SDK example |
|
|
| ```python |
| import os |
| |
| from openai import OpenAI |
| |
| client = OpenAI( |
| base_url="http://localhost:7860/v1", |
| api_key=os.environ["API_PASSWORD"], |
| ) |
| |
| response = client.chat.completions.create( |
| model="model", |
| messages=[{"role": "user", "content": "Hello"}], |
| ) |
| |
| print(response.choices[0].message.content) |
| ``` |
|
|
| ## Hugging Face Spaces notes |
|
|
| This repository is ready for a Docker Space with no frontend. The only required changes at deployment time are the three environment variables above. |
|
|