Neutrino-8B

An 8B chat model whose every transformer linear is stored five-valued (sub-2 bits per weight) in a single 2.56 GB download, and which loads as a native transformers model in one from_pretrained call. One container file runs on CPUs, on Apple silicon through MLX, on NVIDIA GPUs, and through a llama.cpp build.

  • Container: neutrino-8b_v4.bin, 36 layers, hidden 4096, vocabulary 151,936. 3,875,404,812 bytes, sha256 016c6f362ad237a8cf4043601efbc1e766574f76ea58691e78fdef51b4155fa0.
  • 252 packed linears plus int8 embedding tables (untied). The weights stay packed in memory and are decoded inside the matrix kernels.
  • Tokenizer: Qwen3-8B (shipped in this repo).
  • This is a non-thinking model: it ships and is evaluated with thinking disabled (enable_thinking=False).

Architecture

field value
Parameters 8,190,735,360 (6,945,767,424 packed projection + 1,244,659,712 int8 embedding + 308,224 fp32 norm)
Decoder layers 36
Hidden width 4,096
Feed-forward width 12,288, gated (SwiGLU)
Attention grouped-query 4:1, 32 query heads, 8 KV heads, head_dim 128
Rotary embedding full head width (rotary_dims 128), theta 1,000,000
Normalization RMSNorm, eps 1e-6, plus per-head Q/K RMSNorm in attention
Context length 40,960 tokens
Vocabulary 151,936
Embeddings untied, separate int8 input-embedding and output-head tables
KV cache 144 KiB/token fp16 (288 KiB fp32): 0.60 GB at 4k, 4.83 GB at 32k

Files

artifact bytes sha256
neutrino-8b_v4.bin (the container every runtime executes) 3,875,404,812 016c6f362ad237a8cf4043601efbc1e766574f76ea58691e78fdef51b4155fa0
neutrino-8b_v4.tv4z (lossless compressed transport of the same container) 2,559,836,297 83ec7a52bd6733e66fb546e89b9c6aa8b3411af7f648e9c0b2bb56ba50aac4df
gguf/neutrino-8b-fv5.gguf (llama.cpp pack) 4,093,015,136 1c13a34360b2531d28821fb6aee41708f99a040215a2522af37213c6c8c87211

The .tv4z transport decodes back to the container byte for byte; the fermion CLI downloads it and decodes it for you. MANIFEST.json lists the size and sha256 of every file in this repository.

Ways to run it:

  1. pip engine (pip install fermion-research): one command; pulls the container and the platform-matching native runtime from this repo.
  2. Native binaries (bin/): prebuilt CPU runtimes for macOS arm64 and Linux x86-64.
  3. MLX pack (mlx/): Python on Apple silicon, Metal kernels.
  4. GGUF pack + our llama.cpp fork (gguf/): llama.cpp tooling on CPU, CUDA and Metal.
  5. transformers (import fermion): the reference PyTorch path.

Quickstart

pip install fermion-research
printf 'In one sentence, why is the sky blue?\n/exit\n' | \
  fermion chat --model fermionresearch/Neutrino-8B --max-new 48

The first fermion chat downloads about 3.9 GB (container, tokenizer and native runtime); later runs load from the local cache.

Or transformers. Download the repo first: the loader resolves the container relative to a local path, so pass the downloaded directory to from_pretrained:

hf download fermionresearch/Neutrino-8B --local-dir Neutrino-8B \
    --exclude "gguf/*" --exclude "*.tv4z"   # skip the packs other runtimes use
import fermion  # registers the trtc_v4 model type
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Neutrino-8B")   # the downloaded directory
tokenizer = AutoTokenizer.from_pretrained("Neutrino-8B")
inputs = tokenizer.apply_chat_template([{"role": "user", "content": "hi"}],
                                       add_generation_prompt=True,
                                       enable_thinking=False,
                                       return_tensors="pt", return_dict=True)
out = model.generate(**inputs, max_new_tokens=256)  # generation_config: greedy
n = inputs["input_ids"].shape[1]
print(tokenizer.decode(out[0][n:], skip_special_tokens=True))

import fermion must come first: it registers the trtc_v4 model type, and without it AutoTokenizer/AutoModelForCausalLM cannot read this repo's config.json. return_dict=True is required on transformers 4.56 and newer, where apply_chat_template returns a mapping rather than a bare tensor.

If you skip the import you will see this:

ValueError: The checkpoint you are trying to load has model type `trtc_v4`
but Transformers does not recognize this architecture. ... You can update
Transformers with the command `pip install --upgrade transformers`.

Upgrading Transformers does not fix it: trtc_v4 is registered at import time by the fermion-research package, and trust_remote_code=True does not help either, because this repo carries no auto_map. Add import fermion above the Transformers import.

generation_config.json in this repo is greedy with no repetition penalty and no top-k: the setting the published GSM8K, IFEval and BFCL numbers were measured at, so a bare generate(), an lm-eval run and fermion generate all reproduce each other. The conversational and long-form settings are one argument away; see Recommended settings.

GGUF (build our public fork once, branch fermion-fv5 at fermionresearch/llama.cpp, = upstream ggml-org/llama.cpp @ d67c0b41 + the FV5 patch, then standard llama.cpp tooling; full instructions in gguf/README.md):

hf download fermionresearch/Neutrino-8B gguf/neutrino-8b-fv5.gguf --local-dir .
git clone https://github.com/fermionresearch/llama.cpp && cd llama.cpp
git checkout fermion-fv5
cmake -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DGGML_METAL=OFF  # no FV5 Metal kernel yet: CPU/CUDA only
cmake --build build -j --target llama-completion
./build/bin/llama-completion -m ../gguf/neutrino-8b-fv5.gguf \
    -p "Why is the sky blue?" -n 256 --temp 0 -no-cnv

MLX (Apple silicon; run from the mlx/ folder of the downloaded repo, where the fermion_mlx package and its tokenizer files ship; requirements.txt installs mlx itself; full instructions in mlx/README.md):

cd Neutrino-8B/mlx          # the directory downloaded above
pip install -r requirements.txt
python -m fermion_mlx --model ../neutrino-8b_v4.bin --mode chat \
    --tokenizer . --prompt "Why is the sky blue?"

The pip transformers path is the reference path: it keeps the weights packed in memory and matches the native runtime token for token under greedy decoding. For speed, use the native binaries in bin/ (below). Loading and chatting on the CPU path peaks under 8 GiB of resident memory.

Recommended settings, by use case

use case temperature top-p top-k repetition penalty penalty window max new tokens stop
conversation / assistant 0.01 on the C binary · 0 on torch 1.0 not exposed 1.05 256 512 EOS 151645
tool calling / structured output 0 (greedy) — — 1.0 (off) — 256 EOS 151645
long-form prose 0.7 0.95 not exposed 1.05 256 1024 EOS 151645
deterministic / scriptable / benchmarks 0 (greedy) — — 1.0 (off) — 256 EOS 151645
  • Structured output. On tool-shaped prompts the model returns whole, fence-free, parseable JSON 8/8 at every setting tried, including at temperature 0.7. Keep the repetition penalty off for it: JSON repeats ", : and {, and a penalty over that punctuation can break the format.
  • Token budgets. The model is thorough; 512 tokens suit conversation and 1024 suit long-form writing.
  • Top-k is not exposed on any runtime we ship, and this repo's generation_config.json sets none. Top-p is the only shortlist setting.
  • --stop-id 151645 is required on the C binary. Without it the binary generates exactly the token count you asked for and does not stop at EOS.
  • Speculative decoding is exact under greedy decoding (--temperature 0): drafted output is token-identical to plain decoding. With sampling, drafted output follows the same distribution as plain output, but individual tokens differ. The MLX --spec mode is greedy-only. On CUDA the pip path computes in bfloat16, so drafted and plain decoding can pick different tokens at near-ties; use plain decoding there, or check your own pair with fermion verify --model <8B> --draft <0.6B> --device cuda, which exits nonzero when the streams differ.
  • Exactly one repetition penalty applies on the torch path: the fermion CLI sets repetition_penalty=1.0 in its generate() call and installs its own windowed penalty. If you assemble your own call, do the same; two penalties of 1.05 compose to 1.1025.

bin/: prebuilt native runtimes

This repo bundles prebuilt fermion-run binaries next to the weights:

Binary Platform Backend
bin/fermion-run-macos-arm64 macOS arm64 (M1 and later) CPU, NEON dotprod, no dylib dependencies
bin/fermion-run-linux-x64 Linux x86-64 CPU, AVX2, single-file binary (glibc 2.34 or later)

On NVIDIA GPUs, use the pip path with --device cuda or the GGUF pack.

Every binary ships with a <name>.sha256 sidecar (it holds a bare filename, so run the check from inside bin/); hf download writes files 0644, so mark the binary executable before running:

(cd bin && shasum -a 256 -c fermion-run-macos-arm64.sha256)
chmod +x bin/fermion-run-macos-arm64
xattr -d com.apple.quarantine bin/fermion-run-macos-arm64 2>/dev/null || true  # macOS only

pip install fermion-research downloads the model and the platform-matching binary from this repo and uses it for the fast path. The binaries' command line, including session mode and speculative decoding with --draft, is documented in bin/README.md.

Benchmarks

All numbers are measured with thinking disabled.

Meter Score Protocol
MMLU-Redux 67.84 generative, thinking off
IFEval, prompt-strict 73.17 generative, thinking off
IFEval, instruction-strict 80.22 same run, per-instruction grading
IFEval, prompt-loose 76.31 same run, loose extraction
BFCL v3 65.31 macro over 13 subsets, bfcl-eval 2025.10.27.1, thinking off
GSM8K, flexible-extract 51.00 0-shot generative, greedy, 256-token cap
GSM8K, strict/stated format 49.33 same run, answer must appear in the stated format

Reproducing the numbers

The model registers as a native transformers model (import fermion), so the public harnesses run directly. GSM8K with lm-eval:

# first: hf download fermionresearch/Neutrino-8B --local-dir Neutrino-8B  (see Quickstart)
import fermion, lm_eval
from lm_eval.models.huggingface import HFLM
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Neutrino-8B")   # the downloaded directory
tok = AutoTokenizer.from_pretrained("Neutrino-8B")
lm = HFLM(pretrained=model, tokenizer=tok)
lm_eval.simple_evaluate(model=lm, tasks=["gsm8k"], num_fewshot=0)  # flexible-extract

IFEval, MMLU-Redux and BFCL v3 with evalscope 1.4.2 against a vLLM OpenAI-compatible endpoint serving this model, with chat_template_kwargs: {"enable_thinking": false} and bfcl-eval==2025.10.27.1:

pip install evalscope==1.4.2 bfcl-eval==2025.10.27.1
evalscope eval --model Neutrino-8B --api-url http://127.0.0.1:8000/v1 \
    --datasets ifeval mmlu_redux bfcl_v3

Weights: get them and verify them

The container ships in this repository. Fetch it and check the digest; MANIFEST.json carries the same sha256, and fermion info verifies it for you:

hf download fermionresearch/Neutrino-8B neutrino-8b_v4.bin --local-dir .
shasum -a 256 neutrino-8b_v4.bin         # must print 016c6f362...b4155fa0

Speed

Single-stream decode, one container across every row:

Platform Surface Rate Memory
H100 80 GB Fermion CUDA engine, with the Neutrino-0.6B draft 763 tok/s (counting; 402-500 tok/s on other prompt classes) —
H100 80 GB Fermion CUDA engine, plain greedy 396 tok/s —
NVIDIA L4 GGUF pack, full offload 35.4 tok/s 4.68 GiB at 4k context
Apple M5, 16 GB MLX pack 25.0 tok/s 3.67 GiB peak
Apple M5 bin/ native binary, CPU only, 9 threads 24.94 tok/s under 8 GiB resident

On the Linux x86-64 binary, passing a draft container with --draft runs speculative decoding in process: 2.23× over plain decoding on a 16-thread x86-64 CPU (5.95 vs 2.67 tok/s).

Speculative decoding with Neutrino-0.6B

The draft for this 8B is Neutrino-0.6B, a 0.6B container trained to predict this model's greedy choices. Pair them with fermion chat --draft, the MLX pack's --mode spec, the native binaries' --draft, or llama-speculative -md on our llama.cpp fork. Draft and target use the same container format and run in one process.

With greedy decoding, drafted output is token-identical to plain Neutrino-8B decoding: speculation changes the speed, never the text.

prompt class agreement with the 8B end-to-end speedup (H100)
counting 100% ×1.80
facts ~87% ×1.23
prose ~75% ×1.12
chat / explanation ~75% ×1.07
code ~78% ×1.06

The speedup depends on the prompt. The draft is not a chat model; for conversation use Neutrino-0.6B-Chat or this 8B.

Notes

  • The GGUF pack loads through our llama.cpp fork (gguf/); stock llama.cpp, ollama and LM Studio builds do not include the FV5 tensor types.
  • More on the format and the engine: fermionresearch.com/research/.

License and attribution

Weights: Apache-2.0. Based on Qwen/Qwen3-8B (Apache-2.0, Alibaba Cloud); see LICENSE and NOTICE.

bin/ binaries and the compiled kernels in mlx/fermion_mlx/: prebuilt, free to use with these weights; no redistribution outside this repo; no reverse engineering.

Contact: contact@fermionresearch.com

Downloads last month
799
GGUF
Model size
8B params
Architecture
qwen3
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FermionResearch/Neutrino-8B

Finetuned
Qwen/Qwen3-8B
Quantized
(457)
this model
Quantizations
1 model