Instructions to use StandardThinking/StandardOne-3B-SH with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use StandardThinking/StandardOne-3B-SH with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="StandardThinking/StandardOne-3B-SH")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("StandardThinking/StandardOne-3B-SH") model = AutoModelForMultimodalLM.from_pretrained("StandardThinking/StandardOne-3B-SH", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Standard One 3B SH
Updated weights (v2, 2026-10-11). v2 is trained further on hard decision tasks (long policies, trade-offs, judging, dates and numbers). The previous release stays available under the tag
v1.
Version: v2
Standard One 3B SH is Standard One 3B with a joint schema head (SH). It reads the scenario once and
scores every option of every question in a single forward pass, then returns probabilities through the same
POST /v1/systemone contract as Standard One. It does not generate text.
| Backbone | Standard One 3B v2.2 (mistral3, Ministral 3 3B text + Pixtral vision tower), fine-tuned for the head and merged; weights in this repository |
| Head | joint schema head, float32 (head/), reading the final layer and layer 19 |
| Question types | choice (2–255 options), noul (yes/no), score (ordinal scale) |
| Context | up to 32,768 tokens per request |
| Prompt format | the chat template with no default system message (see "Prompt format") |
| Other formats | StandardOne-3B-SH-FP8 · StandardOne-3B-SH-GGUF |
| Larger model | StandardOne-8B-SH |
| License | Apache-2.0 |
Why a schema head
Standard One reads one answer letter per option, so a question is limited by the label alphabet and every question in a request needs its own pass over the scenario. The schema head instead attends over the hidden states of the whole request (the final layer and one intermediate layer), places every question and option in one joint graph, and returns a calibrated distribution for each question at once. In practice:
- Many options. Intent and routing questions with 100–255 options keep nearly the accuracy of small option sets.
- Many questions per scenario. All questions of a request share one pass over the scenario.
- Long scenarios. Decisions buried in long documents are found far more reliably than with Standard One v2.2.
Results
All numbers compare models on identical items, scored the same way.
| Benchmark | Standard One 3B v2.2 | SH v1 | SH v2 (this release) |
|---|---|---|---|
| JevBench public — hard tier (111 items) | 45.95 | 54.05 | 63.96 |
| JevBench public — original tier (72 items) | 93.06 | 97.22 | 98.61 |
| JevBench public — easy tier (48 items) | 100.00 | 100.00 | 100.00 |
| Intent labels with 27–151 options (CLINC150, BANKING77, MASSIVE; 18,000 items) | 79.97 | 92.18 | 91.84 |
| Intent labels with up to 26 options (750 items) | 85.47 | 94.67 | 94.00 |
| Intent labels with 255 options (CLINC150, 1,100 utterances) | 83.80 | 97.02 | 96.70 |
| Decision hidden in a long document (150 items) | 38.00 | 82.00 | 82.00 |
| WinoGrande (dev, 2,010 items) | 80.65 | 92.64 | 92.19 |
| Policy-judgement repair (1,200 items) | 56.67 | 68.08 | 64.58 |
| Game positions: dots and boxes / snake and 2048 / tetris | 37.70 / 59.67 / 55.11 | 49.36 / 64.77 / 65.21 | 48.61 / 65.63 / 64.32 |
| Held-out decision suite (600 items) | 82.00 | 82.67 | 84.33 |
| Held-out realistic decision suite (600 items) | 88.17 | 88.50 | 88.33 |
| Held-out hard decision suite (600 items) | 44.67 | 40.17 | 41.67 |
Accuracy in percent. The Decision Index for v2 is being measured and will be added here.
Changes in v2
- Hard decision tasks. v2 is fine-tuned on about 2,200 new decision items written in the families where v1 was weakest: long policies with exceptions and overrides, multi-criteria trade-offs, judging a response against a rubric, and date and number reasoning. The items were authored by large language models, every gold answer was checked by an independent model that answered without seeing it (and recomputed where they disagreed), and items sharing text with the JevBench public set or our evaluation sets were removed. No JevBench item was used. These items are trained with softened targets (0.8 on the gold option), which keeps the head from becoming over-confident on tasks it cannot always solve.
- Effect. The JevBench hard tier rises from 54.05 to 63.96 (long policies 7 → 13 of 19 items). Confidence on that tier tracks accuracy closely (mean confidence 72.6 for 64.0 correct).
- Trade-offs. Policy-judgement repair is 3.5 points below v1 (64.58 vs 68.08); WinoGrande, games and the intent
label sets are within one point of v1. Keep
v1if policy-judgement repair matters most to you.
Prompt format
The included chat_template.jinja has an empty default system message, and the reference server renders requests
without one (SH_SYSTEM_PROMPT=none, the default in code/render.py). Do not add the Ministral default system
prompt: the head was trained without it, and adding it changes the hidden states it reads.
Quick start
pip install -r code/requirements.txt # torch, transformers, safetensors
hf download StandardThinking/StandardOne-3B-SH --local-dir StandardOne-3B-SH
python StandardOne-3B-SH/code/serve_head.py \
--snapshot StandardOne-3B-SH --head StandardOne-3B-SH/head --port 30171 \
--served-model-name standardthinking/standard-one-3b-sh
curl -s localhost:30171/v1/systemone -H 'content-type: application/json' -d '{
"state": {"message": "Please close my card, I lost it yesterday."},
"questions": {
"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": {"cancel": "close the card", "limit": "change the limit", "other": null}},
"urgent": {"type": "noul", "instructions": "The request is urgent."},
"mood": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]}
}
}'
The server batches concurrent requests (--max-batch-requests, --max-batch-tokens) and binds 127.0.0.1 by
default. It runs on one GPU with the backbone in bfloat16 and the head in float32. On Blackwell GPUs (B200/B300)
install a CUDA 13 build of PyTorch first (pip install torch --index-url https://download.pytorch.org/whl/cu130).
For CPU or Apple-silicon use, see StandardOne-3B-SH-GGUF.
Fast serving with the Standard One engine
The reference server above is plain PyTorch. For production throughput and latency, serve this repository with the Standard One engine, our SGLang fork with a prefill-only decision mode and the schema-head readout (standard-one-sglang, standard-one-adapter, Apache-2.0). It is the engine we run in production. Build the image once (it replaces only the Python package of a stock SGLang image):
mkdir ctx && cd ctx
git clone https://github.com/standardthinkingai/standard-one-sglang sglang-jev
git clone https://github.com/standardthinkingai/standard-one-adapter jev-adapter
mkdir -p jit-cache/aiter jit-cache/root-cache
docker build -f sglang-jev/examples/runtime/jev/rocm/Dockerfile.decision \
--build-arg BASE_IMAGE=lmsysorg/sglang:nightly-dev-cu13-20260929-79cafec0 -t standard-one-engine:cu13 .
Then serve the downloaded repository as it is (weights at the top level, head in head/):
docker run -d --gpus all --network host --ipc host --shm-size 32g \
-v $PWD/StandardOne-3B-SH:/models/m:ro standard-one-engine:cu13 python3 -m sglang.launch_server \
--model-path /models/m --decision-readout-path /models/m/head --served-model-name standardthinking/standard-one-3b-sh \
--host 127.0.0.1 --port 30000 --tp-size 1 --model-impl sglang --dtype bfloat16 --attention-backend triton \
--context-length 32768 --max-running-requests 64 --mem-fraction-static 0.3 \
--chunked-prefill-size -1 --model-config-parser hf --load-format safetensors \
--disable-radix-cache --disable-decode-cuda-graph --disable-prefill-cuda-graph --decision-deadline-seconds 110
It answers POST /v1/systemone on port 30000. Unlike the reference server, the engine requires the model field
(the served model name), as in the System One API:
curl -s localhost:30000/v1/systemone -H 'content-type: application/json' -d '{
"model": "standardthinking/standard-one-3b-sh",
"state": {"message": "Please close my card, I lost it yesterday."},
"questions": {"urgent": {"type": "noul", "instructions": "The request is urgent."}}
}'
Checked on one NVIDIA B300 against the reference
server on 966 evaluation questions: 98.96 % the same answer as the reference server; one small request 24 ms median, one request at a time. On AMD MI350X/MI355X build with
BASE_IMAGE=lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260929 and run with
--device /dev/kfd --device /dev/dri --group-add video -e SGLANG_USE_AITER=0 instead of --gpus all.
Files
| Path | Contents |
|---|---|
model-0000{1,2}-of-00002.safetensors, config.json, tokenizer files |
merged backbone, bfloat16 (Mistral3ForConditionalGeneration) |
chat_template.jinja |
chat template with an empty default system message |
head/ |
schema_head.safetensors (float32) and schema_head_config.json (taps: final layer and layer 19) |
code/ |
reference /v1/systemone server: request rendering, head, scoring |
SHA256SUMS, release-manifest.json |
checksums; source and merge details |
Limitations
- Requests longer than 32,768 tokens are refused, not truncated.
- On our hard held-out decision suite (600 items; long policies, multi-step lookups, date and number reasoning) it scores 41.67 against 44.67 for Standard One 3B v2.2. Prefer Standard One 3B for that kind of workload.
- Text input only has been validated for this release.
- The reference server is a PyTorch server; for throughput use the Standard One engine (see above).
- Probabilities are calibrated on our held-out data; refit thresholds on your own distribution before relying on them.
License
Apache-2.0. Built on Standard One 3B (Apache-2.0), which is built on Ministral 3 3B (Apache-2.0).
- Downloads last month
- -