Instructions to use FermionResearch/Neutrino-8B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use FermionResearch/Neutrino-8B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="FermionResearch/Neutrino-8B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("FermionResearch/Neutrino-8B", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use FermionResearch/Neutrino-8B with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf FermionResearch/Neutrino-8B # Run inference directly in the terminal: llama cli -hf FermionResearch/Neutrino-8B
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf FermionResearch/Neutrino-8B # Run inference directly in the terminal: llama cli -hf FermionResearch/Neutrino-8B
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf FermionResearch/Neutrino-8B # Run inference directly in the terminal: ./llama-cli -hf FermionResearch/Neutrino-8B
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf FermionResearch/Neutrino-8B # Run inference directly in the terminal: ./build/bin/llama-cli -hf FermionResearch/Neutrino-8B
Use Docker
docker model run hf.co/FermionResearch/Neutrino-8B
- LM Studio
- Jan
- vLLM
How to use FermionResearch/Neutrino-8B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "FermionResearch/Neutrino-8B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FermionResearch/Neutrino-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/FermionResearch/Neutrino-8B
- SGLang
How to use FermionResearch/Neutrino-8B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "FermionResearch/Neutrino-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FermionResearch/Neutrino-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "FermionResearch/Neutrino-8B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "FermionResearch/Neutrino-8B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Ollama
How to use FermionResearch/Neutrino-8B with Ollama:
ollama run hf.co/FermionResearch/Neutrino-8B
- Unsloth Desktop
- Pi
How to use FermionResearch/Neutrino-8B with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FermionResearch/Neutrino-8B
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "FermionResearch/Neutrino-8B" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use FermionResearch/Neutrino-8B with Docker Model Runner:
docker model run hf.co/FermionResearch/Neutrino-8B
- Lemonade
How to use FermionResearch/Neutrino-8B with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull FermionResearch/Neutrino-8B
Run and chat with the model
lemonade run user.Neutrino-8B-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use FermionResearch/Neutrino-8B with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FermionResearch/Neutrino-8B
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default FermionResearch/Neutrino-8B
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use FermionResearch/Neutrino-8B with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf FermionResearch/Neutrino-8B
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "FermionResearch/Neutrino-8B" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Neutrino-8B
An 8B chat model whose every transformer linear is stored five-valued
(sub-2 bits per weight) in a single 2.56 GB download, and which loads as
a native transformers model in one from_pretrained call. One container
file runs on CPUs, on Apple silicon through MLX, on NVIDIA GPUs, and through
a llama.cpp build.
- Container:
neutrino-8b_v4.bin, 36 layers, hidden 4096, vocabulary 151,936. 3,875,404,812 bytes, sha256016c6f362ad237a8cf4043601efbc1e766574f76ea58691e78fdef51b4155fa0. - 252 packed linears plus int8 embedding tables (untied). The weights stay packed in memory and are decoded inside the matrix kernels.
- Tokenizer: Qwen3-8B (shipped in this repo).
- This is a non-thinking model: it ships and is evaluated with thinking
disabled (
enable_thinking=False).
Architecture
| field | value |
|---|---|
| Parameters | 8,190,735,360 (6,945,767,424 packed projection + 1,244,659,712 int8 embedding + 308,224 fp32 norm) |
| Decoder layers | 36 |
| Hidden width | 4,096 |
| Feed-forward width | 12,288, gated (SwiGLU) |
| Attention | grouped-query 4:1, 32 query heads, 8 KV heads, head_dim 128 |
| Rotary embedding | full head width (rotary_dims 128), theta 1,000,000 |
| Normalization | RMSNorm, eps 1e-6, plus per-head Q/K RMSNorm in attention |
| Context length | 40,960 tokens |
| Vocabulary | 151,936 |
| Embeddings | untied, separate int8 input-embedding and output-head tables |
| KV cache | 144 KiB/token fp16 (288 KiB fp32): 0.60 GB at 4k, 4.83 GB at 32k |
Files
| artifact | bytes | sha256 |
|---|---|---|
neutrino-8b_v4.bin (the container every runtime executes) |
3,875,404,812 | 016c6f362ad237a8cf4043601efbc1e766574f76ea58691e78fdef51b4155fa0 |
neutrino-8b_v4.tv4z (lossless compressed transport of the same container) |
2,559,836,297 | 83ec7a52bd6733e66fb546e89b9c6aa8b3411af7f648e9c0b2bb56ba50aac4df |
gguf/neutrino-8b-fv5.gguf (llama.cpp pack) |
4,093,015,136 | 1c13a34360b2531d28821fb6aee41708f99a040215a2522af37213c6c8c87211 |
The .tv4z transport decodes back to the container byte for byte; the
fermion CLI downloads it and decodes it for you. MANIFEST.json lists the
size and sha256 of every file in this repository.
Ways to run it:
- pip engine (
pip install fermion-research): one command; pulls the container and the platform-matching native runtime from this repo. - Native binaries (
bin/): prebuilt CPU runtimes for macOS arm64 and Linux x86-64. - MLX pack (
mlx/): Python on Apple silicon, Metal kernels. - GGUF pack + our llama.cpp fork (
gguf/): llama.cpp tooling on CPU, CUDA and Metal. - transformers (
import fermion): the reference PyTorch path.
Quickstart
pip install fermion-research
printf 'In one sentence, why is the sky blue?\n/exit\n' | \
fermion chat --model fermionresearch/Neutrino-8B --max-new 48
The first fermion chat downloads about 3.9 GB (container, tokenizer and
native runtime); later runs load from the local cache.
Or transformers. Download the repo first: the loader resolves the
container relative to a local path, so pass the downloaded directory to
from_pretrained:
hf download fermionresearch/Neutrino-8B --local-dir Neutrino-8B \
--exclude "gguf/*" --exclude "*.tv4z" # skip the packs other runtimes use
import fermion # registers the trtc_v4 model type
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Neutrino-8B") # the downloaded directory
tokenizer = AutoTokenizer.from_pretrained("Neutrino-8B")
inputs = tokenizer.apply_chat_template([{"role": "user", "content": "hi"}],
add_generation_prompt=True,
enable_thinking=False,
return_tensors="pt", return_dict=True)
out = model.generate(**inputs, max_new_tokens=256) # generation_config: greedy
n = inputs["input_ids"].shape[1]
print(tokenizer.decode(out[0][n:], skip_special_tokens=True))
import fermion must come first: it registers the trtc_v4 model type, and
without it AutoTokenizer/AutoModelForCausalLM cannot read this repo's
config.json. return_dict=True is required on transformers 4.56 and
newer, where apply_chat_template returns a mapping rather than a bare
tensor.
If you skip the import you will see this:
ValueError: The checkpoint you are trying to load has model type `trtc_v4`
but Transformers does not recognize this architecture. ... You can update
Transformers with the command `pip install --upgrade transformers`.
Upgrading Transformers does not fix it: trtc_v4 is registered at import
time by the fermion-research package, and trust_remote_code=True does
not help either, because this repo carries no auto_map. Add
import fermion above the Transformers import.
generation_config.json in this repo is greedy with no repetition penalty
and no top-k: the setting the published GSM8K, IFEval and BFCL numbers were
measured at, so a bare generate(), an lm-eval run and fermion generate
all reproduce each other. The conversational and long-form settings are one
argument away; see Recommended settings.
GGUF (build our public fork once, branch fermion-fv5 at
fermionresearch/llama.cpp,
= upstream ggml-org/llama.cpp @ d67c0b41 + the FV5 patch, then
standard llama.cpp tooling; full instructions in gguf/README.md):
hf download fermionresearch/Neutrino-8B gguf/neutrino-8b-fv5.gguf --local-dir .
git clone https://github.com/fermionresearch/llama.cpp && cd llama.cpp
git checkout fermion-fv5
cmake -B build -DCMAKE_BUILD_TYPE=Release -DLLAMA_CURL=OFF -DGGML_METAL=OFF # no FV5 Metal kernel yet: CPU/CUDA only
cmake --build build -j --target llama-completion
./build/bin/llama-completion -m ../gguf/neutrino-8b-fv5.gguf \
-p "Why is the sky blue?" -n 256 --temp 0 -no-cnv
MLX (Apple silicon; run from the mlx/ folder of the downloaded repo, where
the fermion_mlx package and its tokenizer files ship; requirements.txt
installs mlx itself; full instructions in mlx/README.md):
cd Neutrino-8B/mlx # the directory downloaded above
pip install -r requirements.txt
python -m fermion_mlx --model ../neutrino-8b_v4.bin --mode chat \
--tokenizer . --prompt "Why is the sky blue?"
The pip transformers path is the reference path: it keeps the weights
packed in memory and matches the native runtime token for token under
greedy decoding. For speed, use the native binaries in bin/ (below).
Loading and chatting on the CPU path peaks under 8 GiB of resident memory.
Recommended settings, by use case
| use case | temperature | top-p | top-k | repetition penalty | penalty window | max new tokens | stop |
|---|---|---|---|---|---|---|---|
| conversation / assistant | 0.01 on the C binary · 0 on torch | 1.0 | not exposed | 1.05 | 256 | 512 | EOS 151645 |
| tool calling / structured output | 0 (greedy) | — | — | 1.0 (off) | — | 256 | EOS 151645 |
| long-form prose | 0.7 | 0.95 | not exposed | 1.05 | 256 | 1024 | EOS 151645 |
| deterministic / scriptable / benchmarks | 0 (greedy) | — | — | 1.0 (off) | — | 256 | EOS 151645 |
- Structured output. On tool-shaped prompts the model returns whole,
fence-free, parseable JSON 8/8 at every setting tried, including at
temperature 0.7. Keep the repetition penalty off for it: JSON repeats
",:and{, and a penalty over that punctuation can break the format. - Token budgets. The model is thorough; 512 tokens suit conversation and 1024 suit long-form writing.
- Top-k is not exposed on any runtime we ship, and this repo's
generation_config.jsonsets none. Top-p is the only shortlist setting. --stop-id 151645is required on the C binary. Without it the binary generates exactly the token count you asked for and does not stop at EOS.- Speculative decoding is exact under greedy decoding (
--temperature 0): drafted output is token-identical to plain decoding. With sampling, drafted output follows the same distribution as plain output, but individual tokens differ. The MLX--specmode is greedy-only. On CUDA the pip path computes in bfloat16, so drafted and plain decoding can pick different tokens at near-ties; use plain decoding there, or check your own pair withfermion verify --model <8B> --draft <0.6B> --device cuda, which exits nonzero when the streams differ. - Exactly one repetition penalty applies on the torch path: the
fermionCLI setsrepetition_penalty=1.0in itsgenerate()call and installs its own windowed penalty. If you assemble your own call, do the same; two penalties of 1.05 compose to 1.1025.
bin/: prebuilt native runtimes
This repo bundles prebuilt fermion-run binaries next to the weights:
| Binary | Platform | Backend |
|---|---|---|
bin/fermion-run-macos-arm64 |
macOS arm64 (M1 and later) | CPU, NEON dotprod, no dylib dependencies |
bin/fermion-run-linux-x64 |
Linux x86-64 | CPU, AVX2, single-file binary (glibc 2.34 or later) |
On NVIDIA GPUs, use the pip path with --device cuda or the GGUF pack.
Every binary ships with a <name>.sha256 sidecar (it holds a bare
filename, so run the check from inside bin/); hf download writes files
0644, so mark the binary executable before running:
(cd bin && shasum -a 256 -c fermion-run-macos-arm64.sha256)
chmod +x bin/fermion-run-macos-arm64
xattr -d com.apple.quarantine bin/fermion-run-macos-arm64 2>/dev/null || true # macOS only
pip install fermion-research downloads the model and the
platform-matching binary from this repo and uses it for the fast path. The
binaries' command line, including session mode and speculative decoding
with --draft, is documented in bin/README.md.
Benchmarks
All numbers are measured with thinking disabled.
| Meter | Score | Protocol |
|---|---|---|
| MMLU-Redux | 67.84 | generative, thinking off |
| IFEval, prompt-strict | 73.17 | generative, thinking off |
| IFEval, instruction-strict | 80.22 | same run, per-instruction grading |
| IFEval, prompt-loose | 76.31 | same run, loose extraction |
| BFCL v3 | 65.31 | macro over 13 subsets, bfcl-eval 2025.10.27.1, thinking off |
| GSM8K, flexible-extract | 51.00 | 0-shot generative, greedy, 256-token cap |
| GSM8K, strict/stated format | 49.33 | same run, answer must appear in the stated format |
Reproducing the numbers
The model registers as a native transformers model (import fermion), so
the public harnesses run directly. GSM8K with lm-eval:
# first: hf download fermionresearch/Neutrino-8B --local-dir Neutrino-8B (see Quickstart)
import fermion, lm_eval
from lm_eval.models.huggingface import HFLM
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained("Neutrino-8B") # the downloaded directory
tok = AutoTokenizer.from_pretrained("Neutrino-8B")
lm = HFLM(pretrained=model, tokenizer=tok)
lm_eval.simple_evaluate(model=lm, tasks=["gsm8k"], num_fewshot=0) # flexible-extract
IFEval, MMLU-Redux and BFCL v3 with evalscope 1.4.2 against a vLLM
OpenAI-compatible endpoint serving this model, with
chat_template_kwargs: {"enable_thinking": false} and
bfcl-eval==2025.10.27.1:
pip install evalscope==1.4.2 bfcl-eval==2025.10.27.1
evalscope eval --model Neutrino-8B --api-url http://127.0.0.1:8000/v1 \
--datasets ifeval mmlu_redux bfcl_v3
Weights: get them and verify them
The container ships in this repository. Fetch it and check the digest;
MANIFEST.json carries the same sha256, and fermion info verifies it for
you:
hf download fermionresearch/Neutrino-8B neutrino-8b_v4.bin --local-dir .
shasum -a 256 neutrino-8b_v4.bin # must print 016c6f362...b4155fa0
Speed
Single-stream decode, one container across every row:
| Platform | Surface | Rate | Memory |
|---|---|---|---|
| H100 80 GB | Fermion CUDA engine, with the Neutrino-0.6B draft | 763 tok/s (counting; 402-500 tok/s on other prompt classes) | — |
| H100 80 GB | Fermion CUDA engine, plain greedy | 396 tok/s | — |
| NVIDIA L4 | GGUF pack, full offload | 35.4 tok/s | 4.68 GiB at 4k context |
| Apple M5, 16 GB | MLX pack | 25.0 tok/s | 3.67 GiB peak |
| Apple M5 | bin/ native binary, CPU only, 9 threads |
24.94 tok/s | under 8 GiB resident |
On the Linux x86-64 binary, passing a draft container with --draft runs
speculative decoding in process: 2.23× over plain decoding on a 16-thread
x86-64 CPU (5.95 vs 2.67 tok/s).
Speculative decoding with Neutrino-0.6B
The draft for this 8B is
Neutrino-0.6B, a
0.6B container trained to predict this model's greedy choices. Pair them
with fermion chat --draft, the MLX pack's --mode spec, the native
binaries' --draft, or llama-speculative -md on our llama.cpp fork. Draft
and target use the same container format and run in one process.
With greedy decoding, drafted output is token-identical to plain Neutrino-8B decoding: speculation changes the speed, never the text.
| prompt class | agreement with the 8B | end-to-end speedup (H100) |
|---|---|---|
| counting | 100% | ×1.80 |
| facts | ~87% | ×1.23 |
| prose | ~75% | ×1.12 |
| chat / explanation | ~75% | ×1.07 |
| code | ~78% | ×1.06 |
The speedup depends on the prompt. The draft is not a chat model; for conversation use Neutrino-0.6B-Chat or this 8B.
Notes
- The GGUF pack loads through our llama.cpp fork (
gguf/); stock llama.cpp, ollama and LM Studio builds do not include the FV5 tensor types. - More on the format and the engine: fermionresearch.com/research/.
License and attribution
Weights: Apache-2.0. Based on Qwen/Qwen3-8B (Apache-2.0, Alibaba Cloud);
see LICENSE and NOTICE.
bin/ binaries and the compiled kernels in mlx/fermion_mlx/: prebuilt,
free to use with these weights; no redistribution outside this repo; no
reverse engineering.
Contact: contact@fermionresearch.com
- Downloads last month
- 799
We're not able to determine the quantization variants.