Instructions to use StandardThinking/StandardOne-3B-SH-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use StandardThinking/StandardOne-3B-SH-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M # Run inference directly in the terminal: llama cli -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M # Run inference directly in the terminal: ./llama-cli -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Use Docker
docker model run hf.co/StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
- LM Studio
- Jan
- Ollama
How to use StandardThinking/StandardOne-3B-SH-GGUF with Ollama:
ollama run hf.co/StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
- Unsloth Desktop
- Pi
How to use StandardThinking/StandardOne-3B-SH-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use StandardThinking/StandardOne-3B-SH-GGUF with Docker Model Runner:
docker model run hf.co/StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
- Lemonade
How to use StandardThinking/StandardOne-3B-SH-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Run and chat with the model
lemonade run user.StandardOne-3B-SH-GGUF-Q4_K_M
List all available models
lemonade list
- Hermes Agent
How to use StandardThinking/StandardOne-3B-SH-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use StandardThinking/StandardOne-3B-SH-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "StandardThinking/StandardOne-3B-SH-GGUF:Q4_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Standard One 3B SH — GGUF
Updated weights (v2, 2026-10-11). This repository now holds GGUF builds of Standard One 3B SH v2, in eight quant levels. The previous builds stay available under the tag
v1.
Version: v2 (2026-10-11)
GGUF builds of the backbone of Standard One 3B SH v2
(mistral3 architecture, text model) for use with llama.cpp, together
with the schema head and a small runner that serves the same POST /v1/systemone contract on CPU, Apple silicon or
any other llama.cpp backend.
The schema head does not read text output. It reads two hidden states for every token — the final layer (after the
last norm) and the output of decoder layer 19 — so stock llama-server cannot run it on its own. The runner here
takes both states from llama.cpp through its evaluation callback (runner/sh_states.cpp) and runs the head in
PyTorch.
Files
| File | Quant | Size |
|---|---|---|
StandardOne-3B-SH-BF16.gguf |
BF16 (no quantization) | 6.9 GB |
StandardOne-3B-SH-Q8_0.gguf |
Q8_0 | 3.7 GB |
StandardOne-3B-SH-Q6_K.gguf |
Q6_K | 2.8 GB |
StandardOne-3B-SH-Q5_K_M.gguf |
Q5_K_M | 2.5 GB |
StandardOne-3B-SH-Q5_K_S.gguf |
Q5_K_S | 2.4 GB |
StandardOne-3B-SH-Q4_K_M.gguf |
Q4_K_M | 2.1 GB |
StandardOne-3B-SH-Q4_K_S.gguf |
Q4_K_S | 2.1 GB |
StandardOne-3B-SH-Q4_0.gguf |
Q4_0 | 2.0 GB |
head/ |
schema head, float32 (69.5M parameters) and its config | 278 MB |
tokenizer/ |
tokenizer and chat template used to render requests (empty default system message) | 34 MB |
runner/ |
sh_states.cpp (state extractor), serve_head_gguf.py (server) and the shared request/head code |
— |
All quant levels are plain llama-quantize builds (no importance matrix). SHA256 checksums: SHA256SUMS. Source
and conversion details: release-manifest.json.
Quick start
# 1. llama.cpp (tested at commit 53ed051) and the state extractor
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp && git checkout 53ed051
cmake -B build && cmake --build build -j
hf download StandardThinking/StandardOne-3B-SH-GGUF StandardOne-3B-SH-Q8_0.gguf \
--include "head/*" "tokenizer/*" "runner/*" --local-dir ../sh-gguf
g++ -O2 -std=c++17 -Iinclude -Iggml/include ../sh-gguf/runner/sh_states.cpp \
-Lbuild/bin -lllama -lggml -lggml-base -Wl,-rpath,$PWD/build/bin -o ../sh-gguf/runner/sh-states
cd ../sh-gguf
# 2. the server (torch CPU build is enough)
pip install torch transformers safetensors numpy tokenizers
python runner/serve_head_gguf.py --gguf StandardOne-3B-SH-Q8_0.gguf --head head --tokenizer tokenizer \
--sh-states runner/sh-states --port 30171 --threads 16
Requests and responses are the same as for StandardOne-3B-SH.
The runner serves one request at a time (one llama.cpp context); --context sets the longest request in tokens
(default 34,000).
With a CUDA build of llama.cpp (cmake -B build -DGGML_CUDA=ON), start the server with GGML_CUDA_DISABLE_GRAPHS=1
in its environment. On an NVIDIA B300 we saw one request hang inside a CUDA-graph launch with graphs enabled; with
them disabled every check below ran without a stall.
Validation
Measured 2026-10-11 through the runner on one NVIDIA B300 (CUDA build of llama.cpp 53ed051, CUDA graphs disabled), against the BF16 reference server of StandardOne-3B-SH v2 on the same 1,431 questions: the three JevBench public tiers, our 600-item hard calibration set and our 600-item held-out decision set.
| Build | JevBench easy (48) | JevBench original (72) | JevBench hard (111) | Hard calibration (600) | Held-out (600) | Same answer as the BF16 server |
|---|---|---|---|---|---|---|
| BF16 server (reference) | 100.00 | 98.61 | 63.96 | 42.33 | 84.33 | — |
| BF16 GGUF | 100.00 | 98.61 | 62.16 | 42.67 | 84.17 | 99.44 % (1,423/1,431) |
| Q8_0 | 100.00 | 98.61 | 62.16 | 42.67 | 83.67 | 99.02 % (1,417/1,431) |
| Q6_K | 100.00 | 98.61 | 60.36 | 42.00 | 83.00 | 97.62 % (1,397/1,431) |
| Q5_K_M | 100.00 | 98.61 | 62.16 | 42.17 | 82.67 | 95.74 % (1,370/1,431) |
| Q5_K_S | 100.00 | 98.61 | 62.16 | 43.00 | 82.67 | 95.32 % (1,364/1,431) |
| Q4_K_M | 100.00 | 98.61 | 63.96 | 42.00 | 82.67 | 93.01 % (1,331/1,431) |
| Q4_K_S | 100.00 | 98.61 | 62.16 | 42.17 | 83.50 | 93.15 % (1,333/1,431) |
| Q4_0 | 100.00 | 100.00 | 64.86 | 39.67 | 83.33 | 92.52 % (1,324/1,431) |
Q8_0 is the recommended build. Below Q6_K, accuracy on these sets stays close but more individual answers move.
Limitations
- Text input only (no vision projector in this repository).
- CPU prefill is slow for long requests: prefill runs on the CPU, so expect several seconds per thousand tokens on a 16-core server.
- The runner is a reference implementation, not an inference engine: one request at a time, no batching.
License
Apache-2.0, the same terms as StandardOne-3B-SH and Standard One 3B, built on Ministral 3 3B (Apache-2.0).
- Downloads last month
- -
4-bit
5-bit
6-bit
8-bit
16-bit
Model tree for StandardThinking/StandardOne-3B-SH-GGUF
Base model
mistralai/Ministral-3-3B-Base-2512