Standard One 8B SH — GGUF

Version: v2 (2026-10-09)

GGUF builds of the backbone of Standard One 8B SH v2 (mistral3 architecture, text model) for use with llama.cpp, together with the schema head and a small runner that serves the same POST /v1/systemone contract on CPU, Apple silicon or any other llama.cpp backend.

The schema head does not read text output. It reads two hidden states for every token — the final layer (after the last norm) and the output of decoder layer 25 — so stock llama-server cannot run it on its own. The runner here takes both states from llama.cpp through its evaluation callback (runner/sh_states.cpp) and runs the head in PyTorch.

Files

File Quant Size
StandardOne-8B-SH-BF16.gguf BF16 (no quantization) 17.0 GB
StandardOne-8B-SH-Q8_0.gguf Q8_0 9.0 GB
StandardOne-8B-SH-Q5_K_M.gguf Q5_K_M 6.1 GB
StandardOne-8B-SH-Q4_K_M.gguf Q4_K_M 5.2 GB
head/ schema head, float32 (70.6M parameters) and its config 282 MB
tokenizer/ tokenizer and chat template used to render requests (empty default system message) 33 MB
runner/ sh_states.cpp (state extractor), serve_head_gguf.py (server) and the shared request/head code —

All quant levels are plain llama-quantize builds (no importance matrix). SHA256 checksums: SHA256SUMS. Source and conversion details: release-manifest.json.

Quick start

# 1. llama.cpp (tested at commit 53ed051) and the state extractor
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp && git checkout 53ed051
cmake -B build && cmake --build build -j
hf download StandardThinking/StandardOne-8B-SH-GGUF StandardOne-8B-SH-Q8_0.gguf \
  --include "head/*" "tokenizer/*" "runner/*" --local-dir ../sh-gguf
g++ -O2 -std=c++17 -Iinclude -Iggml/include ../sh-gguf/runner/sh_states.cpp \
  -Lbuild/bin -lllama -lggml -lggml-base -Wl,-rpath,$PWD/build/bin -o ../sh-gguf/runner/sh-states
cd ../sh-gguf

# 2. the server (torch CPU build is enough)
pip install torch transformers safetensors numpy tokenizers
python runner/serve_head_gguf.py --gguf StandardOne-8B-SH-Q8_0.gguf --head head --tokenizer tokenizer \
  --sh-states runner/sh-states --port 30171 --threads 16

Requests and responses are the same as for StandardOne-8B-SH. The runner serves one request at a time (one llama.cpp context); --context sets the longest request in tokens (default 34,000).

Validation

Measured 2026-10-09 on CPU (no GPU offload) through the runner, against the BF16 reference server of StandardOne-8B-SH v2 on the same items. Realistic and hard-proxy rows are fixed subsets (every 10th / every 7th item) of the 600-item held-out suites.

Build JevBench easy (48) JevBench original (72) JevBench hard (111) Realistic (60) Hard proxy (80) Same answer as the BF16 server
BF16 server (reference) 100.00 98.61 63.96 93.33 48.75 —
BF16 GGUF 100.00 98.61 64.86 93.33 48.75 99.73 % (370/371)
Q8_0 100.00 98.61 63.96 91.67 47.50 99.19 % (368/371)
Q5_K_M 100.00 98.61 63.96 91.67 50.00 97.30 % (361/371)
Q4_K_M 100.00 97.22 64.86 91.67 48.75 95.69 % (355/371)

The hidden states themselves match closely: on a sample request the BF16 GGUF states have mean cosine 0.99999 with the transformers states (final layer and layer 25), Q8_0 0.9996–0.9997, Q4_K_M 0.984–0.989. Q8_0 is the recommended build; Q4_K_M is the smallest one that stays usable.

Limitations

  • Text input only (no vision projector in this repository).
  • CPU prefill is slow for long requests: Q8_0 reads about 140 tokens per second on 16 server cores, so a 3,000-token request takes about 20 seconds.
  • The runner is a reference implementation, not an inference engine: one request at a time, no batching.

License

Apache-2.0, the same terms as StandardOne-8B-SH and Standard One 8B, built on Ministral 3 8B (Apache-2.0).

Downloads last month
-
GGUF
Model size
8B params
Architecture
mistral3
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for StandardThinking/StandardOne-8B-SH-GGUF