Instructions to use StandardThinking/StandardOne-8B-SH-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use StandardThinking/StandardOne-8B-SH-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Use Docker
docker model run hf.co/StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use StandardThinking/StandardOne-8B-SH-GGUF with Ollama:
ollama run hf.co/StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use StandardThinking/StandardOne-8B-SH-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use StandardThinking/StandardOne-8B-SH-GGUF with Docker Model Runner:
docker model run hf.co/StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
- Lemonade
How to use StandardThinking/StandardOne-8B-SH-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.StandardOne-8B-SH-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use StandardThinking/StandardOne-8B-SH-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use StandardThinking/StandardOne-8B-SH-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "StandardThinking/StandardOne-8B-SH-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Standard One 8B SH — GGUF
Version: v2 (2026-10-09)
GGUF builds of the backbone of Standard One 8B SH v2
(mistral3 architecture, text model) for use with llama.cpp, together
with the schema head and a small runner that serves the same POST /v1/systemone contract on CPU, Apple silicon or
any other llama.cpp backend.
The schema head does not read text output. It reads two hidden states for every token — the final layer (after the
last norm) and the output of decoder layer 25 — so stock llama-server cannot run it on its own. The runner here
takes both states from llama.cpp through its evaluation callback (runner/sh_states.cpp) and runs the head in
PyTorch.
Files
| File | Quant | Size |
|---|---|---|
StandardOne-8B-SH-BF16.gguf |
BF16 (no quantization) | 17.0 GB |
StandardOne-8B-SH-Q8_0.gguf |
Q8_0 | 9.0 GB |
StandardOne-8B-SH-Q5_K_M.gguf |
Q5_K_M | 6.1 GB |
StandardOne-8B-SH-Q4_K_M.gguf |
Q4_K_M | 5.2 GB |
head/ |
schema head, float32 (70.6M parameters) and its config | 282 MB |
tokenizer/ |
tokenizer and chat template used to render requests (empty default system message) | 33 MB |
runner/ |
sh_states.cpp (state extractor), serve_head_gguf.py (server) and the shared request/head code |
— |
All quant levels are plain llama-quantize builds (no importance matrix). SHA256 checksums: SHA256SUMS. Source
and conversion details: release-manifest.json.
Quick start
# 1. llama.cpp (tested at commit 53ed051) and the state extractor
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp && git checkout 53ed051
cmake -B build && cmake --build build -j
hf download StandardThinking/StandardOne-8B-SH-GGUF StandardOne-8B-SH-Q8_0.gguf \
--include "head/*" "tokenizer/*" "runner/*" --local-dir ../sh-gguf
g++ -O2 -std=c++17 -Iinclude -Iggml/include ../sh-gguf/runner/sh_states.cpp \
-Lbuild/bin -lllama -lggml -lggml-base -Wl,-rpath,$PWD/build/bin -o ../sh-gguf/runner/sh-states
cd ../sh-gguf
# 2. the server (torch CPU build is enough)
pip install torch transformers safetensors numpy tokenizers
python runner/serve_head_gguf.py --gguf StandardOne-8B-SH-Q8_0.gguf --head head --tokenizer tokenizer \
--sh-states runner/sh-states --port 30171 --threads 16
Requests and responses are the same as for StandardOne-8B-SH.
The runner serves one request at a time (one llama.cpp context); --context sets the longest request in tokens
(default 34,000).
Validation
Measured 2026-10-09 on CPU (no GPU offload) through the runner, against the BF16 reference server of StandardOne-8B-SH v2 on the same items. Realistic and hard-proxy rows are fixed subsets (every 10th / every 7th item) of the 600-item held-out suites.
| Build | JevBench easy (48) | JevBench original (72) | JevBench hard (111) | Realistic (60) | Hard proxy (80) | Same answer as the BF16 server |
|---|---|---|---|---|---|---|
| BF16 server (reference) | 100.00 | 98.61 | 63.96 | 93.33 | 48.75 | — |
| BF16 GGUF | 100.00 | 98.61 | 64.86 | 93.33 | 48.75 | 99.73 % (370/371) |
| Q8_0 | 100.00 | 98.61 | 63.96 | 91.67 | 47.50 | 99.19 % (368/371) |
| Q5_K_M | 100.00 | 98.61 | 63.96 | 91.67 | 50.00 | 97.30 % (361/371) |
| Q4_K_M | 100.00 | 97.22 | 64.86 | 91.67 | 48.75 | 95.69 % (355/371) |
The hidden states themselves match closely: on a sample request the BF16 GGUF states have mean cosine 0.99999 with the transformers states (final layer and layer 25), Q8_0 0.9996–0.9997, Q4_K_M 0.984–0.989. Q8_0 is the recommended build; Q4_K_M is the smallest one that stays usable.
Limitations
- Text input only (no vision projector in this repository).
- CPU prefill is slow for long requests: Q8_0 reads about 140 tokens per second on 16 server cores, so a 3,000-token request takes about 20 seconds.
- The runner is a reference implementation, not an inference engine: one request at a time, no batching.
License
Apache-2.0, the same terms as StandardOne-8B-SH and Standard One 8B, built on Ministral 3 8B (Apache-2.0).
- Downloads last month
- -
4-bit
5-bit
8-bit
16-bit
Model tree for StandardThinking/StandardOne-8B-SH-GGUF
Base model
mistralai/Ministral-3-8B-Base-2512