Standard One 3B SH — GGUF

Updated weights (v2, 2026-10-11). This repository now holds GGUF builds of Standard One 3B SH v2, in eight quant levels. The previous builds stay available under the tag v1.

Version: v2 (2026-10-11)

GGUF builds of the backbone of Standard One 3B SH v2 (mistral3 architecture, text model) for use with llama.cpp, together with the schema head and a small runner that serves the same POST /v1/systemone contract on CPU, Apple silicon or any other llama.cpp backend.

The schema head does not read text output. It reads two hidden states for every token — the final layer (after the last norm) and the output of decoder layer 19 — so stock llama-server cannot run it on its own. The runner here takes both states from llama.cpp through its evaluation callback (runner/sh_states.cpp) and runs the head in PyTorch.

Files

File Quant Size
StandardOne-3B-SH-BF16.gguf BF16 (no quantization) 6.9 GB
StandardOne-3B-SH-Q8_0.gguf Q8_0 3.7 GB
StandardOne-3B-SH-Q6_K.gguf Q6_K 2.8 GB
StandardOne-3B-SH-Q5_K_M.gguf Q5_K_M 2.5 GB
StandardOne-3B-SH-Q5_K_S.gguf Q5_K_S 2.4 GB
StandardOne-3B-SH-Q4_K_M.gguf Q4_K_M 2.1 GB
StandardOne-3B-SH-Q4_K_S.gguf Q4_K_S 2.1 GB
StandardOne-3B-SH-Q4_0.gguf Q4_0 2.0 GB
head/ schema head, float32 (69.5M parameters) and its config 278 MB
tokenizer/ tokenizer and chat template used to render requests (empty default system message) 34 MB
runner/ sh_states.cpp (state extractor), serve_head_gguf.py (server) and the shared request/head code —

All quant levels are plain llama-quantize builds (no importance matrix). SHA256 checksums: SHA256SUMS. Source and conversion details: release-manifest.json.

Quick start

# 1. llama.cpp (tested at commit 53ed051) and the state extractor
git clone https://github.com/ggml-org/llama.cpp && cd llama.cpp && git checkout 53ed051
cmake -B build && cmake --build build -j
hf download StandardThinking/StandardOne-3B-SH-GGUF StandardOne-3B-SH-Q8_0.gguf \
  --include "head/*" "tokenizer/*" "runner/*" --local-dir ../sh-gguf
g++ -O2 -std=c++17 -Iinclude -Iggml/include ../sh-gguf/runner/sh_states.cpp \
  -Lbuild/bin -lllama -lggml -lggml-base -Wl,-rpath,$PWD/build/bin -o ../sh-gguf/runner/sh-states
cd ../sh-gguf

# 2. the server (torch CPU build is enough)
pip install torch transformers safetensors numpy tokenizers
python runner/serve_head_gguf.py --gguf StandardOne-3B-SH-Q8_0.gguf --head head --tokenizer tokenizer \
  --sh-states runner/sh-states --port 30171 --threads 16

Requests and responses are the same as for StandardOne-3B-SH. The runner serves one request at a time (one llama.cpp context); --context sets the longest request in tokens (default 34,000).

With a CUDA build of llama.cpp (cmake -B build -DGGML_CUDA=ON), start the server with GGML_CUDA_DISABLE_GRAPHS=1 in its environment. On an NVIDIA B300 we saw one request hang inside a CUDA-graph launch with graphs enabled; with them disabled every check below ran without a stall.

Validation

Measured 2026-10-11 through the runner on one NVIDIA B300 (CUDA build of llama.cpp 53ed051, CUDA graphs disabled), against the BF16 reference server of StandardOne-3B-SH v2 on the same 1,431 questions: the three JevBench public tiers, our 600-item hard calibration set and our 600-item held-out decision set.

Build JevBench easy (48) JevBench original (72) JevBench hard (111) Hard calibration (600) Held-out (600) Same answer as the BF16 server
BF16 server (reference) 100.00 98.61 63.96 42.33 84.33 —
BF16 GGUF 100.00 98.61 62.16 42.67 84.17 99.44 % (1,423/1,431)
Q8_0 100.00 98.61 62.16 42.67 83.67 99.02 % (1,417/1,431)
Q6_K 100.00 98.61 60.36 42.00 83.00 97.62 % (1,397/1,431)
Q5_K_M 100.00 98.61 62.16 42.17 82.67 95.74 % (1,370/1,431)
Q5_K_S 100.00 98.61 62.16 43.00 82.67 95.32 % (1,364/1,431)
Q4_K_M 100.00 98.61 63.96 42.00 82.67 93.01 % (1,331/1,431)
Q4_K_S 100.00 98.61 62.16 42.17 83.50 93.15 % (1,333/1,431)
Q4_0 100.00 100.00 64.86 39.67 83.33 92.52 % (1,324/1,431)

Q8_0 is the recommended build. Below Q6_K, accuracy on these sets stays close but more individual answers move.

Limitations

  • Text input only (no vision projector in this repository).
  • CPU prefill is slow for long requests: prefill runs on the CPU, so expect several seconds per thousand tokens on a 16-core server.
  • The runner is a reference implementation, not an inference engine: one request at a time, no batching.

License

Apache-2.0, the same terms as StandardOne-3B-SH and Standard One 3B, built on Ministral 3 3B (Apache-2.0).

Downloads last month
-
GGUF
Model size
3B params
Architecture
mistral3
Hardware compatibility
Log In to add your hardware

4-bit

5-bit

6-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for StandardThinking/StandardOne-3B-SH-GGUF