Instructions to use StandardThinking/StandardOne-8B-SH-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use StandardThinking/StandardOne-8B-SH-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="StandardThinking/StandardOne-8B-SH-FP8")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("StandardThinking/StandardOne-8B-SH-FP8") model = AutoModelForMultimodalLM.from_pretrained("StandardThinking/StandardOne-8B-SH-FP8", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Standard One 8B SH — FP8
Version: v2 (2026-10-09)
An FP8 (compressed-tensors, float8_e4m3 weights, dynamic per-token activations) quantization of the merged backbone
of Standard One 8B SH v2, with the same schema head.
The language-model linear projections (q/k/v/o, gate/up/down; 238 tensors) are quantized per output channel
(scale = absmax / 448); the vision tower, multi-modal projector, embeddings, norms and lm_head stay in BF16. This is
the same data-free recipe and quantization_config as
StandardOne-8B-FP8. The head (head/) is float32 and
identical to the BF16 release.
| Weights | ~10 GB (BF16 release: 17 GB) |
| Head | head/schema_head.safetensors, float32, taps: final layer and layer 25 |
| Prompt format | chat template with no default system message (see the main card) |
Validation
Both precisions were served the same way by our SGLang-based decision engine (schema-head readout, bfloat16 activations, one option order), measured 2026-10-09. Accuracy is the most probable answer.
| Suite | BF16 (StandardOne-8B-SH v2) | FP8 (this repository) | Change (points) |
|---|---|---|---|
| many-option questions, 53–151 options (18,000) | 92.46 % | 92.52 % | +0.06 |
| the same question set, at most 26 options (750) | 95.33 % | 95.33 % | +0.00 |
| long-document questions (150) | 86.00 % | 86.00 % | +0.00 |
| held-out decision set (600) | 90.50 % | 91.00 % | +0.50 |
| hard proxy (600) | 45.67 % | 46.00 % | +0.33 |
| realistic transfer set (600) | 91.00 % | 90.67 % | −0.33 |
| WinoGrande dev (2,010) | 95.12 % | 94.98 % | −0.14 |
| policy-judgement repair (1,200) | 73.58 % | 73.50 % | −0.08 |
| JevBench public easy (48) | 100.00 % | 100.00 % | +0.00 |
| JevBench public standard (72) | 98.61 % | 98.61 % | +0.00 |
| JevBench public hard (111) | 63.96 % | 67.57 % | +3.61 |
Across 4,791 validation questions, FP8 and BF16 gave the same answer for 97.89 %.
Notes
- Engines that load compressed-tensors FP8 (SGLang, vLLM) read this checkpoint. On ROCm with the native per-token activation-quantization fallback, SGLang 0.5.20 needs a small fix when the engine pads the token dimension (the per-token scale buffer is padded but the fallback writes unpadded rows); without it the server fails at warmup. The validation above used that fix.
- The PyTorch reference server in StandardOne-8B-SH is written for the BF16 checkpoint; use the BF16 release with it.
License
Apache-2.0, the same terms as StandardOne-8B-SH, built on Standard One 8B and Ministral 3 8B (Apache-2.0).
- Downloads last month
- -
Model tree for StandardThinking/StandardOne-8B-SH-FP8
Base model
mistralai/Ministral-3-8B-Base-2512