Instructions to use StandardThinking/StandardOne-3B-SH-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use StandardThinking/StandardOne-3B-SH-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="StandardThinking/StandardOne-3B-SH-FP8")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("StandardThinking/StandardOne-3B-SH-FP8") model = AutoModelForMultimodalLM.from_pretrained("StandardThinking/StandardOne-3B-SH-FP8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Standard One 3B SH — FP8
Updated weights (v2, 2026-10-11). This repository now holds the FP8 build of Standard One 3B SH v2. The previous build stays available under the tag
v1.
Version: v2 (2026-10-11)
An FP8 (compressed-tensors, float8_e4m3 weights, dynamic per-token activations) quantization of the merged backbone
of Standard One 3B SH v2, with the same schema head.
The language-model linear projections (q/k/v/o, gate/up/down; 182 tensors) are quantized per output channel
(scale = absmax / 448); the vision tower, multi-modal projector, embeddings, norms and lm_head stay in BF16. This is
the same data-free recipe and quantization_config as
StandardOne-3B-FP8. The head (head/) is float32 and
identical to the BF16 release.
| Weights | 4.4 GB (BF16 release: 7.7 GB) |
| Head | head/schema_head.safetensors, float32, taps: final layer and layer 19 |
| Prompt format | chat template with no default system message (see the main card) |
Serving
Serve it with the Standard One engine (our SGLang fork with the schema-head readout; build and run commands in the
main card).
Use this repository as --model-path and its head/ as --decision-readout-path; the other arguments are the same.
Engines that load compressed-tensors FP8 read the backbone; the schema head needs the Standard One engine or the
PyTorch reference server. The reference server in StandardOne-3B-SH is written for the BF16 checkpoint; use the BF16
release with it.
Validation
Both precisions were served the same way by the Standard One engine on one NVIDIA B300 (schema-head readout, bfloat16 activations, one option order), measured 2026-10-11. Accuracy is the most probable answer.
| Suite | BF16 (StandardOne-3B-SH v2) | FP8 (this repository) | Change (points) |
|---|---|---|---|
| many-option questions, 53–151 options (18,000) | 91.83 % | 91.86 % | +0.03 |
| the same question set, at most 26 options (750) | 93.87 % | 94.27 % | +0.40 |
| long-document questions (150) | 80.67 % | 83.33 % | +2.67 |
| held-out decision set (600) | 84.33 % | 84.17 % | −0.17 |
| hard proxy (600) | 41.00 % | 39.33 % | −1.67 |
| realistic transfer set (600) | 88.50 % | 88.33 % | −0.17 |
| WinoGrande dev (2,010) | 92.14 % | 91.79 % | −0.35 |
| policy-judgement repair (1,200) | 64.33 % | 64.50 % | +0.17 |
| game positions: dots and boxes / snake and 2048 / tetris | 48.79 / 65.68 / 64.38 % | 49.09 / 66.64 / 64.19 % | +0.31 / +0.96 / −0.19 |
| JevBench public easy (48) | 100.00 % | 100.00 % | +0.00 |
| JevBench public standard (72) | 98.61 % | 100.00 % | +1.39 |
| JevBench public hard (111) | 62.16 % | 61.26 % | −0.90 |
Across 11,428 validation questions, FP8 and BF16 gave the same answer for 96.77 %.
Notes
- On AMD ROCm with the native per-token activation-quantization fallback, SGLang 0.5.20 needs a small fix when the engine pads the token dimension (the per-token scale buffer is padded but the fallback writes unpadded rows); without it the server fails at warmup. This release was validated on NVIDIA (CUDA), where the fallback is not used.
License
Apache-2.0, the same terms as StandardOne-3B-SH, built on Standard One 3B and Ministral 3 3B (Apache-2.0).
- Downloads last month
- -
Model tree for StandardThinking/StandardOne-3B-SH-FP8
Base model
mistralai/Ministral-3-3B-Base-2512