Standard One 8B SH — FP8

Version: v2 (2026-10-09)

An FP8 (compressed-tensors, float8_e4m3 weights, dynamic per-token activations) quantization of the merged backbone of Standard One 8B SH v2, with the same schema head. The language-model linear projections (q/k/v/o, gate/up/down; 238 tensors) are quantized per output channel (scale = absmax / 448); the vision tower, multi-modal projector, embeddings, norms and lm_head stay in BF16. This is the same data-free recipe and quantization_config as StandardOne-8B-FP8. The head (head/) is float32 and identical to the BF16 release.

Weights ~10 GB (BF16 release: 17 GB)
Head head/schema_head.safetensors, float32, taps: final layer and layer 25
Prompt format chat template with no default system message (see the main card)

Validation

Both precisions were served the same way by our SGLang-based decision engine (schema-head readout, bfloat16 activations, one option order), measured 2026-10-09. Accuracy is the most probable answer.

Suite BF16 (StandardOne-8B-SH v2) FP8 (this repository) Change (points)
many-option questions, 53–151 options (18,000) 92.46 % 92.52 % +0.06
the same question set, at most 26 options (750) 95.33 % 95.33 % +0.00
long-document questions (150) 86.00 % 86.00 % +0.00
held-out decision set (600) 90.50 % 91.00 % +0.50
hard proxy (600) 45.67 % 46.00 % +0.33
realistic transfer set (600) 91.00 % 90.67 % −0.33
WinoGrande dev (2,010) 95.12 % 94.98 % −0.14
policy-judgement repair (1,200) 73.58 % 73.50 % −0.08
JevBench public easy (48) 100.00 % 100.00 % +0.00
JevBench public standard (72) 98.61 % 98.61 % +0.00
JevBench public hard (111) 63.96 % 67.57 % +3.61

Across 4,791 validation questions, FP8 and BF16 gave the same answer for 97.89 %.

Notes

  • Engines that load compressed-tensors FP8 (SGLang, vLLM) read this checkpoint. On ROCm with the native per-token activation-quantization fallback, SGLang 0.5.20 needs a small fix when the engine pads the token dimension (the per-token scale buffer is padded but the fallback writes unpadded rows); without it the server fails at warmup. The validation above used that fix.
  • The PyTorch reference server in StandardOne-8B-SH is written for the BF16 checkpoint; use the BF16 release with it.

License

Apache-2.0, the same terms as StandardOne-8B-SH, built on Standard One 8B and Ministral 3 8B (Apache-2.0).

Downloads last month
-
Safetensors
Model size
9B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for StandardThinking/StandardOne-8B-SH-FP8