Standard One 3B SH — FP8

Updated weights (v2, 2026-10-11). This repository now holds the FP8 build of Standard One 3B SH v2. The previous build stays available under the tag v1.

Version: v2 (2026-10-11)

An FP8 (compressed-tensors, float8_e4m3 weights, dynamic per-token activations) quantization of the merged backbone of Standard One 3B SH v2, with the same schema head. The language-model linear projections (q/k/v/o, gate/up/down; 182 tensors) are quantized per output channel (scale = absmax / 448); the vision tower, multi-modal projector, embeddings, norms and lm_head stay in BF16. This is the same data-free recipe and quantization_config as StandardOne-3B-FP8. The head (head/) is float32 and identical to the BF16 release.

Weights 4.4 GB (BF16 release: 7.7 GB)
Head head/schema_head.safetensors, float32, taps: final layer and layer 19
Prompt format chat template with no default system message (see the main card)

Serving

Serve it with the Standard One engine (our SGLang fork with the schema-head readout; build and run commands in the main card). Use this repository as --model-path and its head/ as --decision-readout-path; the other arguments are the same. Engines that load compressed-tensors FP8 read the backbone; the schema head needs the Standard One engine or the PyTorch reference server. The reference server in StandardOne-3B-SH is written for the BF16 checkpoint; use the BF16 release with it.

Validation

Both precisions were served the same way by the Standard One engine on one NVIDIA B300 (schema-head readout, bfloat16 activations, one option order), measured 2026-10-11. Accuracy is the most probable answer.

Suite BF16 (StandardOne-3B-SH v2) FP8 (this repository) Change (points)
many-option questions, 53–151 options (18,000) 91.83 % 91.86 % +0.03
the same question set, at most 26 options (750) 93.87 % 94.27 % +0.40
long-document questions (150) 80.67 % 83.33 % +2.67
held-out decision set (600) 84.33 % 84.17 % −0.17
hard proxy (600) 41.00 % 39.33 % −1.67
realistic transfer set (600) 88.50 % 88.33 % −0.17
WinoGrande dev (2,010) 92.14 % 91.79 % −0.35
policy-judgement repair (1,200) 64.33 % 64.50 % +0.17
game positions: dots and boxes / snake and 2048 / tetris 48.79 / 65.68 / 64.38 % 49.09 / 66.64 / 64.19 % +0.31 / +0.96 / −0.19
JevBench public easy (48) 100.00 % 100.00 % +0.00
JevBench public standard (72) 98.61 % 100.00 % +1.39
JevBench public hard (111) 62.16 % 61.26 % −0.90

Across 11,428 validation questions, FP8 and BF16 gave the same answer for 96.77 %.

Notes

  • On AMD ROCm with the native per-token activation-quantization fallback, SGLang 0.5.20 needs a small fix when the engine pads the token dimension (the per-token scale buffer is padded but the fallback writes unpadded rows); without it the server fails at warmup. This release was validated on NVIDIA (CUDA), where the fallback is not used.

License

Apache-2.0, the same terms as StandardOne-3B-SH, built on Standard One 3B and Ministral 3 3B (Apache-2.0).

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for StandardThinking/StandardOne-3B-SH-FP8