StandardOne-8B-SH / README.md
MyeongHoJeong's picture
v3: further training on hard decision tasks (merged weights + head)
6b952a1 verified
|
Raw History Blame Contribute Delete
11.2 kB
---
base_model: StandardThinking/StandardOne-8B
library_name: transformers
license: apache-2.0
pipeline_tag: text-classification
language:
- en
tags:
- mistral3
- decision-model
- typed-decisions
- schema-head
- jev
- calibration
- decode-free
---
# Standard One 8B SH
> **Updated weights (v3, 2026-10-11).** v3 is trained further on hard decision tasks (long policies, trade-offs,
> judging, dates and numbers). Earlier releases stay available under the tags `v2` (merged weights) and `v1`
> (LoRA adapter).
**Version:** v3
**Standard One 8B SH** is Standard One 8B with a **joint schema head** (SH). It reads the scenario once and
scores every option of every question in a single forward pass, then returns probabilities through the same
`POST /v1/systemone` contract as Standard One. It does not generate text.
| | |
|---|---|
| Backbone | Standard One 8B v2.2 (`mistral3`, Ministral 3 8B text + Pixtral vision tower), fine-tuned for the head and merged; weights in this repository |
| Head | joint schema head, 70.6M parameters, float32 (`head/`), reading the final layer and layer 25 |
| Question types | `choice` (2–255 options), `noul` (yes/no), `score` (ordinal scale) |
| Context | up to 32,768 tokens per request |
| Prompt format | the chat template with **no default system message** (see "Prompt format") |
| Other formats | [StandardOne-8B-SH-FP8](https://huggingface.co/StandardThinking/StandardOne-8B-SH-FP8) · [StandardOne-8B-SH-GGUF](https://huggingface.co/StandardThinking/StandardOne-8B-SH-GGUF) (both still hold v2) |
| Smaller model | [StandardOne-3B-SH](https://huggingface.co/StandardThinking/StandardOne-3B-SH) |
| License | Apache-2.0 |
## Why a schema head
Standard One reads one answer letter per option, so a question is limited by the label alphabet and every
question in a request needs its own pass over the scenario. The schema head instead attends over the hidden states
of the whole request (the final layer and one intermediate layer), places every question and option in one joint
graph, and returns a calibrated distribution for each question at once. In practice:
- **Many options.** Intent and routing questions with 100–255 options keep nearly the accuracy of small option sets.
- **Many questions per scenario.** All questions of a request share one pass over the scenario.
- **Long scenarios.** Decisions buried in long documents are found far more reliably than with Standard One v2.2.
## Results
All numbers compare models on identical items, scored the same way.
| Benchmark | Standard One 8B v2.2 | SH v2 | **SH v3 (this release)** |
|---|---:|---:|---:|
| JevBench public — hard tier (111 items) | 56.76 | 63.96 | **74.77** |
| JevBench public — original tier (72 items) | **98.61** | **98.61** | 95.83 |
| JevBench public — easy tier (48 items) | 100.00 | 100.00 | 100.00 |
| Intent labels with 27–151 options (CLINC150, BANKING77, MASSIVE; 18,000 items) | 82.67 | **92.50** | 92.48 |
| Intent labels with up to 26 options (750 items) | 88.13 | 95.20 | **95.47** |
| Intent labels with 255 options (CLINC150, 1,100 utterances) | 85.34 | **97.43** | 97.32 |
| Decision hidden in a long document (150 items) | 40.00 | 84.67 | **87.33** |
| WinoGrande (dev, 2,010 items) | 87.86 | **95.07** | 94.98 |
| Policy-judgement repair (1,200 items) | 60.58 | **72.92** | 72.50 |
| Game positions: dots and boxes / snake and 2048 / tetris | 38.37 / 58.63 / 54.86 | **50.42** / 67.59 / **65.90** | 49.05 / **68.00** / 63.81 |
| Held-out decision suite (600 items) | 82.50 | **90.67** | 87.67 |
| Held-out realistic decision suite (600 items) | **91.50** | 90.83 | 89.67 |
| Held-out hard decision suite (600 items) | **53.50** | 46.00 | 41.17 |
Accuracy in percent. The v3 column was measured with the Standard One engine (see "Fast serving"), which gives the same
answer as the reference server in `code/` on 99.58 % of the 1,431 questions we checked (all 231 JevBench public answers
identical); the other columns with the reference server. v2 scored 48.68 on the Decision Index 0.2.1 (see tag `v2`);
the Decision Index for v3 is being measured and will be added here.
## Changes in v3
- **Hard decision tasks.** v3 is fine-tuned on about 2,200 new decision items written in the families where v2 was
weakest: long policies with exceptions and overrides, multi-criteria trade-offs, judging a response against a
rubric, and date and number reasoning. The items were authored by large language models, every gold answer was
checked by an independent model that answered without seeing it (and recomputed where they disagreed), and items
sharing text with the JevBench public set or our evaluation sets were removed. No JevBench item was used. These
items are trained with softened targets (0.8 on the gold option), which keeps the head from becoming over-confident
on tasks it cannot always solve. A short final pass on policy, long-document and multi-step items keeps those
skills from regressing.
- **Effect.** The JevBench hard tier rises from 63.96 to 74.77. Confidence on that tier tracks accuracy closely (mean
confidence 76.8 for 74.8 correct; expected calibration error 0.045, v2 0.097). Long-document decisions rise from
84.67 to 87.33.
- **Trade-offs.** The held-out decision suite is 3.0 points below v2 (87.67 vs 90.67), the hard held-out suite 4.8
points (41.17 vs 46.00) and the JevBench original tier two items (95.83 vs 98.61). Intent labels, WinoGrande and
policy judgement are within half a point of v2. Keep `v2` if those held-out suites matter most to you.
## Prompt format
The included `chat_template.jinja` has an empty default system message, and the reference server renders requests
without one (`SH_SYSTEM_PROMPT=none`, the default in `code/render.py`). Do not add the Ministral default system
prompt: the head was trained without it, and adding it changes the hidden states it reads.
## Quick start
```bash
pip install -r code/requirements.txt # torch, transformers, safetensors
hf download StandardThinking/StandardOne-8B-SH --local-dir StandardOne-8B-SH
python StandardOne-8B-SH/code/serve_head.py \
--snapshot StandardOne-8B-SH --head StandardOne-8B-SH/head --port 30171
```
```bash
curl -s localhost:30171/v1/systemone -H 'content-type: application/json' -d '{
"state": {"message": "Please close my card, I lost it yesterday."},
"questions": {
"intent": {"type": "choice", "instructions": "What does the customer want?",
"criteria": {"cancel": "close the card", "limit": "change the limit", "other": null}},
"urgent": {"type": "noul", "instructions": "The request is urgent."},
"mood": {"type": "score", "instructions": "How upset is the customer?", "criteria": ["calm", "annoyed", "angry"]}
}
}'
```
The server batches concurrent requests (`--max-batch-requests`, `--max-batch-tokens`) and binds `127.0.0.1` by
default. It runs on one GPU with the backbone in bfloat16 and the head in float32. On Blackwell GPUs (B200/B300)
install a CUDA 13 build of PyTorch first (`pip install torch --index-url https://download.pytorch.org/whl/cu130`); the
CUDA 12.8 build lacks kernels for those GPUs.
## Fast serving with the Standard One engine
The reference server above is plain PyTorch. For production throughput and latency, serve this repository with the
**Standard One engine**, our SGLang fork with a prefill-only decision mode and the schema-head readout
([standard-one-sglang](https://github.com/standardthinkingai/standard-one-sglang),
[standard-one-adapter](https://github.com/standardthinkingai/standard-one-adapter), Apache-2.0). It is the engine we run
in production. Build the image once (it replaces only the Python package of a stock SGLang image):
```bash
mkdir ctx && cd ctx
git clone https://github.com/standardthinkingai/standard-one-sglang sglang-jev
git clone https://github.com/standardthinkingai/standard-one-adapter jev-adapter
mkdir -p jit-cache/aiter jit-cache/root-cache
docker build -f sglang-jev/examples/runtime/jev/rocm/Dockerfile.decision \
--build-arg BASE_IMAGE=lmsysorg/sglang:nightly-dev-cu13-20260929-79cafec0 -t standard-one-engine:cu13 .
```
Then serve the downloaded repository as it is (weights at the top level, head in `head/`):
```bash
docker run -d --gpus all --network host --ipc host --shm-size 32g \
-v $PWD/StandardOne-8B-SH:/models/m:ro standard-one-engine:cu13 python3 -m sglang.launch_server \
--model-path /models/m --decision-readout-path /models/m/head --served-model-name standardthinking/standard-one-8b-sh \
--host 127.0.0.1 --port 30000 --tp-size 1 --model-impl sglang --dtype bfloat16 --attention-backend triton \
--context-length 32768 --max-running-requests 64 --mem-fraction-static 0.3 \
--chunked-prefill-size -1 --model-config-parser hf --load-format safetensors \
--disable-radix-cache --disable-decode-cuda-graph --disable-prefill-cuda-graph --decision-deadline-seconds 110
```
It answers `POST /v1/systemone` on port 30000. Unlike the reference server, the engine requires the `model` field
(the served model name), as in the System One API:
```bash
curl -s localhost:30000/v1/systemone -H 'content-type: application/json' -d '{
"model": "standardthinking/standard-one-8b-sh",
"state": {"message": "Please close my card, I lost it yesterday."},
"questions": {"urgent": {"type": "noul", "instructions": "The request is urgent."}}
}'
```
Checked on one NVIDIA B300 against the reference server on 1,431 evaluation questions: 99.58 % the same answer, all
231 JevBench public answers identical. On AMD MI350X/MI355X build with
`BASE_IMAGE=lmsysorg/sglang-rocm:v0.5.20-rocm720-mi35x-20260929` and run with
`--device /dev/kfd --device /dev/dri --group-add video -e SGLANG_USE_AITER=0` instead of `--gpus all`.
## Files
| Path | Contents |
|---|---|
| `model-0000{1..4}-of-00004.safetensors`, `config.json`, tokenizer files | merged backbone, bfloat16 (`Mistral3ForConditionalGeneration`) |
| `chat_template.jinja` | chat template with an empty default system message |
| `head/` | `schema_head.safetensors` (float32) and `schema_head_config.json` (taps: final layer and layer 25) |
| `code/` | reference `/v1/systemone` server: request rendering, head, scoring |
| `SHA256SUMS`, `release-manifest.json` | checksums; source and merge details |
## Limitations
- Requests longer than 32,768 tokens are refused, not truncated. This is a serving limit: the backbone handles longer
inputs, but the head was not trained on them.
- On our hard held-out decision suite (600 items; long policies, multi-step lookups, date and number reasoning)
it scores 41.17 against 53.50 for Standard One 8B v2.2. Prefer Standard One 8B for that kind of workload.
- Text input only has been validated for this release.
- The reference server is a PyTorch server; for throughput use the Standard One engine (see above).
- Probabilities are calibrated on our held-out data; refit thresholds on your own distribution before relying on them.
## License
Apache-2.0. Built on [Standard One 8B](https://huggingface.co/StandardThinking/StandardOne-8B) (Apache-2.0), which is
built on [Ministral 3 8B](https://huggingface.co/mistralai/Ministral-3-8B-Instruct-2512-BF16) (Apache-2.0).