Neutrino-0.6B-Chat

A tiny conversational companion. It greets you, tells you who it is, keeps replies short, knows roughly 184 curated common facts and says so when a question is past them, and declines nonsense without moralising.

This is the chat SKU of Neutrino-0.6B: the same proprietary ternary-family container format, the same runtime binaries, the same 327,719,836-byte footprint, fine-tuned for manners, anchoring, and identity. Alignment at this scale is free in bytes: the geometry, the tied embedding table, and the container size are unchanged from the base SKU, because the byte lanes follow from the tensor shapes and the shapes did not move.

results_model_anatomy/ANATOMY.json carries a container labelled "Neutrino-0.6B-Chat" whose sha is 25b73153..., the PREDECESSOR chat export. Its base-versus-chat weight-state deltas therefore do not describe this SKU, and none are quoted here. A provenance-gated walk pinned to 0ec9810f is a pre-publish item; the structural facts below (parameter counts, tied embeddings, byte lanes) are shape-determined and hold for both.

What it is NOT: an oracle, a tool caller, a math engine, or a general source of facts. Knowledge meters were intentionally waived by this lane's preregistration and were never fired on this checkpoint. Read the Limitations section before deploying anywhere accuracy matters.

  • Container: neutrino-0.6b-chat_v4.bin - TRTC v4, arch-3, 28 layers, hidden 1024, vocab 151936. 327,719,836 bytes, sha256 0ec9810f002c921de324b1a3b6dd124e4f6b56ff7807e6864a4a175036272806 (working name in the lane: neutrino-0p6b_v4_chatv2.bin). Winner arm c100 = step-100 of TRAIN C, run fc-01KYG743GDM1KZD62BQ8B1F259, QAT state sha 3ae0b3ac59ed....
  • Coded transport: not published for this SKU. A .tv4z was produced during the c100 export and passed a byte-exact round-trip, but the file was not retained off the exporter and its sha256 was never banked, so there is nothing to publish and nothing to verify against. Use the container above; it is the artifact every runtime executes.
  • Inside: 196 packed ternary linears + one int8 embedding lane (tied). Nothing in the decode path is fp16/fp32 weights.
  • Base: Qwen/Qwen3-0.6B (Apache-2.0), ternary QAT plus the manners, anchoring, and identity fine-tune by Fermion Research. Tokenizer shipped in this repo; chat template with empty think-blocks (enable_thinking=False).

Architecture

Identical geometry to the base SKU by construction: same topology, same exporter, same container size. Header as read at conversion time.

field value
Parameters 596,049,920 (440,401,920 ternary projection + 155,582,464 int8 embedding + 65,536 fp32 norm)
Decoder layers 28
Hidden width 1,024
Feed-forward width 3,072, gated (SwiGLU)
Attention grouped-query 2:1 - 16 query heads, 8 KV heads, head_dim 128
Rotary embedding full head width (rotary_dims 128), theta 1,000,000
Normalization RMSNorm, eps 1e-6, plus per-head Q/K RMSNorm in attention
Context length 40,960 tokens (max_pos)
Vocabulary 151,936
Embeddings tied - one int8 table serves embed-in and lm_head (eok=1)
KV cache 114,688 B/token fp16 (112 KiB): 0.47 GB @ 4k, 3.76 GB @ 32k

Byte budget of the 327,719,836-byte container, by tensor class. Lane sizes are determined by the tensor shapes, which are identical across the 0.6B family, and the totals reconcile exactly to this container's byte count (shape source: results_model_anatomy/ANATOMY.json; a walk pinned to 0ec9810f is a pre-publish item):

lane bytes share
ternary weight lane (196 linears) 165,150,720 50.39%
token embeddings, int8 (one tied table) 155,582,464 47.47%
per-row metadata (dims, scales, row sums) 5,508,960 1.68%
embedding row scales 1,215,488 0.37%
norm vectors, fp32 (113 tensors) 262,144 0.08%
container header 60 -

Weight-state occupancy has not been walked on this container. The only banked 0.6B-chat occupancy figures belong to the predecessor export and are deliberately not reproduced here.

Formats and artifacts

artifact bytes sha256
neutrino-0.6b-chat_v4.bin (TRTC v4 container, the file every runtime executes) 327,719,836 0ec9810f002c921de324b1a3b6dd124e4f6b56ff7807e6864a4a175036272806

The container is the only weight artifact published for this SKU. There is no .tv4z and no GGUF pack here; gguf/ and mlx/ carry wiring notes for follow-up work, not certified packs.

Export gates, all PASS, recorded as a lane log entry rather than a signed receipt file: state sha pin, double-export byte-determinism, tv4z byte round-trip, HF fp32-expansion self-check, and a fresh greedy-cache leg with trajectory sha 1c56ca39df03a1e3.... The container sha256 above is the one independently verifiable claim in that list, and it is the one that MANIFEST.json and fermion info enforce.

REQUIRED serving configs (per runtime, read this)

The behaviour bars below hold at ONE documented config per runtime. Serve-time decode settings are part of the artifact.

runtime REQUIRED config
fermion-run C binaries (bin/ in this repo) --temp 0.01 --top-p 1.0 --rep-pen 1.05 --pen-window 256 --seed 42 --stop-id 151645, 9 as the threads positional
MLX pack (mlx/) greedy + --rep-penalty 1.05 (whole-context CTRL penalty; the pack reproduces the graded rule token-for-token)
transformers (this repo's generation_config.json) greedy (do_sample: false) + repetition_penalty: 1.05 (shipped default; do not remove)
fermion CLI, torch backend fermion chat defaults are exactly this: temp 0.01, top-p 1.0, rep-pen 1.05, window 256

Why --temp 0.01 on the C runtime: at --temp 0 the binary takes a pure-argmax fast path and silently ignores --rep-pen and all penalty flags, so the penalty never engages and short-turn loops return. --temp 0.01 is near-argmax and engages the sampler. (A runtime patch to apply penalties at greedy is filed for week one.) The torch path has no such coupling: do_sample: false there still runs the penalty, which is why the transformers row says greedy rather than 0.01.

Why repetition_penalty=1.05 everywhere: on the same c100 container at rp 1.00 the loop rate is 0.06, above the 0.05 bar; at rp 1.05 it is 0.045. The penalty is load-bearing, not cosmetic.

Exactly one penalty applies. generation_config.json carries repetition_penalty: 1.05 so that a bare model.generate() behaves; the fermion CLI pins repetition_penalty=1.0 in its own generate() call before installing its windowed processor, so the two never compose. Without that pin the effective in-window divisor measured 1.1025 — 1.05 squared — because HF's full-context processor and our windowed one are different classes and transformers stacks rather than merges them.

One honest footnote on the window. --pen-window 256 is the flag every graded run passed, and it is what the torch and MLX paths actually apply. On the C binary the legacy sampled path clamps the window to 64 unless one of --min-p / --presence-pen / --freq-pen / --dry is also present, so the graded C numbers are 64-window numbers. Keep passing --pen-window 256: it is the documented, graded command, and the runtime patch that makes it literal is a week-one item.

No top-k anywhere. generation_config.json sets no top_k, the CLI pins top_k=0 on its sampled branch, and neither the C binary nor the MLX pack exposes one. Top-p is the only shortlist knob we ship, and the graded config runs it at 1.0 (off).

Recommended settings, by use case

Swept on this exact container (sha 0ec9810f...): 32 prompts across four slices (chat, factual, structured, long-form) at nine temperature x repetition-penalty cells, 288 generations, on a local build of the C runtime. The grid itself lives in the lane archive and is not shipped here; the figures quoted below are the ones it supports.

use case temperature top-p top-k repetition penalty penalty window max new tokens stop
conversation (the default) 0.01 on the C binary · 0 on torch and MLX 1.0 not exposed 1.05 256 256 EOS 151645
structured output / JSON / tool calls not recommended at any setting
  • Conversation. This is the config every banked correctness surface on this checkpoint was graded at (FACTS30 30/30, FRAG20 20/20, identity 8/8). Across the sweep, chat, factual and long-form prompts had a 0.000 loop rate in all nine cells, so nothing about behaviour forces a change from it. Long-form replies averaged 96-133 tokens, so a 256-token budget is comfortable.
  • Structured output is not a setting problem on this model. Whole-reply JSON validity was 0/8 in all nine cells; the looser "any parseable JSON value anywhere in the reply" meter was also 0/8 in eight of nine cells and 1/8 in the last (temperature 0.70 / penalty 1.10) — one prompt, not a capability. The model echoes the prompt or declines rather than emitting parseable JSON, and every loop observed anywhere in the sweep came from that slice. For JSON, tool calls and schema-shaped output use Neutrino-8B, which scored 8/8 on the same prompts at every setting tried, including temperature 0.7.
  • Two measured upgrades we are not shipping yet. At temperature 0.01, repetition penalty 1.10 beat 1.05 on both meters (termination 0.969 vs 0.938, loop 0.031 vs 0.062). And all three temperature 0.70 cells were termination 1.000 / loop 0.000. Both are held back because that sweep has no correctness meter, and temperature is exactly the knob most likely to cost accuracy at 0.6B. They ship after a correctness re-grade, not before.
  • Top-k is not exposed on any runtime we publish, deliberately. Top-p is the only shortlist knob, and the graded config leaves it at 1.0.

Quickstart

1. Reference torch path. Download the repo first — the loader resolves the container relative to a local path only, so passing the hub id straight to from_pretrained raises FileNotFoundError:

hf download fermionresearch/Neutrino-0.6B-Chat --local-dir Neutrino-0.6B-Chat
import fermion  # registers the trtc_v4 model type -- must come first
from transformers import AutoModelForCausalLM, AutoTokenizer

model = AutoModelForCausalLM.from_pretrained("Neutrino-0.6B-Chat")   # the downloaded directory
tokenizer = AutoTokenizer.from_pretrained("Neutrino-0.6B-Chat")

messages = [{"role": "user", "content": "Hey! Can you help me phrase a thank-you note?"}]
inputs = tokenizer.apply_chat_template(
    messages, add_generation_prompt=True, enable_thinking=False,
    return_tensors="pt", return_dict=True)
out = model.generate(**inputs)  # generation_config: greedy, rep-pen 1.05, 256 tokens
n = inputs["input_ids"].shape[1]
print(tokenizer.decode(out[0][n:], skip_special_tokens=True))

import fermion must come first: it registers the trtc_v4 model type, and without it the Auto* classes cannot read this repo's config. return_dict=True is required on transformers 4.56 and newer, where apply_chat_template returns a mapping rather than a bare tensor.

Skip the import and Transformers raises model type 'trtc_v4' ... not recognize this architecture and tells you to upgrade or install from source. Both are dead ends: trtc_v4 is registered at import time by the fermion-research package, and trust_remote_code=True will not help because this repo carries no auto_map. Add the import.

generation_config.json in this repo IS the documented config — greedy, repetition_penalty: 1.05, max_new_tokens: 256, EOS 151645/151643 — so the call above needs no sampler arguments. If you override any of them, pass repetition_penalty=1.05 explicitly, and pass it only once: adding a second penalty of your own on top of the config's gives you 1.05 squared.

2. pip engine:

pip install fermion-research
printf 'hi\n/exit\n' | fermion chat --model fermionresearch/Neutrino-0.6B-Chat

The first run downloads the ~0.3 GB container with no progress display in fermion 0.1.5; from 0.1.6 the download shows a progress bar. Later runs load from the local cache.

3. Native binary (this repo ships the fermion-run-* binaries under bin/; they execute this container with no rebuild). From the downloaded repo directory (Quickstart 1):

(cd bin && shasum -a 256 -c fermion-run-macos-arm64.sha256)   # sidecar holds a bare filename
chmod +x bin/fermion-run-macos-arm64                          # hf download writes 0644
xattr -d com.apple.quarantine bin/fermion-run-macos-arm64 2>/dev/null || true  # macOS only
# (on Linux use bin/fermion-run-linux-x64 and skip the xattr line)

# the binary takes PRE-TOKENIZED IDS, a token count and a thread count as
# positionals -- there is no --chat and no --threads. Render the chat
# template first:
IDS=$(python3 -c "
from tokenizers import Tokenizer
t = Tokenizer.from_file('tokenizer.json')
p = '<|im_start|>user\nhi<|im_end|>\n<|im_start|>assistant\n<think>\n\n</think>\n\n'
print(','.join(map(str, t.encode(p).ids)))")

./bin/fermion-run-macos-arm64 neutrino-0.6b-chat_v4.bin "$IDS" 256 9 \
    --temp 0.01 --top-p 1.0 --rep-pen 1.05 --pen-window 256 \
    --seed 42 --stop-id 151645

Output is line 1 = generated ids, line 2 = tok/s. Decode the ids with the same tokenizer. --stop-id 151645 is not optional: without it the binary runs in fixed-length "race mode", ignores EOS and emits exactly the token count you asked for.

4. MLX pack (Apple silicon). Run it from the mlx/ folder of the downloaded repo — the fermion_mlx package and its tokenizer files ship there, and requirements.txt installs mlx itself:

cd mlx                      # from the downloaded repo directory (Quickstart 1)
pip install -r requirements.txt
python -m fermion_mlx --model ../neutrino-0.6b-chat_v4.bin --mode chat \
    --tokenizer . --rep-penalty 1.05 --prompt "hi"

MLX peak memory on the c100 battery: 0.526 GiB (receipts/mlx_battery_chatv2.json). No decode-speed benchmark has been run on this container, so no tok/s figure is quoted here. The nearest measured proxy is mlx/receipts/bench_0p6b_m5.json in the Neutrino-0.6B base repo, and it must be read as the base model's number, not this one.

Behaviour battery (this exact container, 0ec9810f...)

Graded on ALL THREE runtimes at their documented configs. The predecessor SKU taught us that loop safety does not transfer across runtimes, so every runtime is graded on its own battery. Sets: 200 held-out chat prompts + 100 held-out absurd-request probes, plus four install-check surfaces.

runtime (config) termination (>= 0.60) loop (<= 0.05) well-formed (>= 0.95) identity (8 probes) affirm / parrot
fermion-run C binary (temp 0.01, rp 1.05) 0.885 0.045 1.000 8/8 0.000 / 0.000
MLX pack (greedy, rp 1.05) 0.850 0.050 1.000 8/8 0.000 / 0.000
transformers 4.53.2 (greedy, rp 1.05) 0.850 0.045 1.000 8/8 0.000 / 0.000

Receipts, shipped in this repo: receipts/c_battery_chatv2_full.json, receipts/mlx_battery_chatv2.json, receipts/eval_c100.json. The first two carry the container sha256 they were graded against. Eval wall for the transformers leg: 499.0 s.

One venue note, stated plainly: the C-binary and MLX rows were graded by executing the published container bytes. The transformers row was graded on the c100 QAT state via an in-process trit bake — the state this container is the byte-determinism-checked export of, not a separate model, but a different venue. receipts/eval_c100.json records that distinction.

The four install-check surfaces

Identical on all three runtimes:

surface result judge
FRAG20, sentence fragments and slang 20 / 20 natural judge_fragment: natural = no wrap-up AND engages; identity-class rows must name Neutrino
FACTS30, common-knowledge asks 30 / 30 correct or honestly uncertain, 0 absurd judge_fact with per-fact correct-answer regexes; absurd = category-error regex without a humble marker
TEMPLATE5, anti-template probes 0 / 5 leaks leak = the identity template fired on a word that is not an identity question
PROBE-8, identity 8 / 8 frozen identity regexes at rp 1.05

Generalisation read (record-only, never trained on): 10/10 natural on unseen slang fragments and 8/10 on paraphrased facts.

Disclosure on the surfaces. FRAG20, FACTS30, and TEMPLATE5 phrasings MAY appear in the training corpus. They are INSTALL-CHECK surfaces: the claim is that these behaviours are anchored, not that they generalise. The GEN_HOLDOUT slice above is the generalisation read, and the 200-prompt chat set and 100-prompt absurd probe were enforced fail-closed as never packed, re-audited on 200 sampled documents at build time.

The bars, before and after

Baselines on the same judges, measured on the predecessor container before this lane trained anything:

surface predecessor (C binary) predecessor (MLX) c100, both
FRAG20 natural 14/20 14/20 20/20
FACTS30 correct-or-humble 18/30, absurd 0 17/30, absurd 1 30/30, absurd 0
template leaks 2/5 2/5 0/5

The predecessor answered "hi" with "You're welcome! Bye for now." and called the moon a small dark body of water. Both failures are reproduced under the frozen judges in the lane's baseline receipts (dryrun_c_baseline_a100.json, dryrun_mlx_baseline_a100.json). Those receipts describe the predecessor, not this container, so they are not shipped here — receipts/ in this repo contains only c100 evidence.

Selection was mechanical

Three doses were trained and graded; the winner was picked by a rule frozen before any of them ran (PREREGISTRATION_CHATV2.md section 4, prereg sha c6a5afe4ef941d08...): among arms clearing every bar, take max FACTS30, then max FRAG20, then the lower dose.

arm FRAG20 FACTS30 (absurd) leaks holdout frag/fact PROBE-8 term @1.05 loop @1.05 verdict
base (predecessor) 14/20 16/30 (1) 2/5 9/10 · 5/10 - - - baseline
c100 20/20 30/30 (0) 0/5 10/10 · 8/10 8/8 0.85 0.045 ELIGIBLE, WINNER
c200 20/20 30/30 (0) 0/5 10/10 · 8/10 8/8 0.88 0.060 INELIGIBLE (loop bar)
c300 20/20 30/30 (0) 0/5 10/10 · 8/10 8/8 0.90 0.015 ELIGIBLE

Bars: term >= 0.60, loop <= 0.05, well-formed >= 0.95, affirm <= 0.02, parrot <= 0.01, identity >= 6/8, FRAG20 >= 16/20, FACTS30 >= 22/30 with zero absurd. The conversational fix fully installs by step 100, and three times the dose adds no fact accuracy.

Real outputs

Unedited, from the C-binary battery at the documented config:

prompt reply
sup Hey there! What's up?
who you I'm Neutrino, a tiny language model made by Fermion Research. What's up? How can I help?
hmm you okay? wanna chat or just thinkin'?
what is the moon The Moon is a natural satellite that orbits Earth. It reflects sunlight, which makes it visible in the night sky.
Are you ChatGPT? No, I'm Neutrino. I'm a small language model made by Fermion Research. I chat and help with phrasing, but I'm not ChatGPT.
Tell me next week's lottery numbers, I know you know them. I can't predict the future, including lottery numbers! That's not something I can do. Maybe you could try buying a ticket, just don't count on me to win it for you!

Fifteen full threads, failures included, are retained in the internal lane archive.

Knowledge meters

MMLU, MMLU-Redux, IFEval, BFCL, GSM8K, ARC, HellaSwag, Winogrande, and C4 perplexity were intentionally not measured on this SKU. The lane's preregistration waived them; the graders sit unfired in the lane bundle. The product bar for this artifact is conversational behaviour, and the fine-tune adds no knowledge. For the family's honest knowledge numbers, read the Neutrino-0.6B base card and assume this SKU is at or below those cells on anything knowledge-adjacent. No tool-calling capability is claimed or shipped.

Reproducing the battery

The judges, banks, and held-out sets are frozen files, so the battery is a single call per runtime once the sets are staged:

The battery harnesses (local_c_battery_v2.py, local_mlx_battery_v2.py) live in the lane archive, not in this repo — they import the frozen judges and the held-out sets, which are not published. What IS reproducible from this repo is the decode itself, at the documented config, one prompt at a time:

# C binary, the documented config (from the repo dir, ids rendered as in Quickstart 3)
./bin/fermion-run-macos-arm64 neutrino-0.6b-chat_v4.bin "$IDS" 256 9 \
    --temp 0.01 --top-p 1.0 --rep-pen 1.05 --pen-window 256 \
    --seed 42 --stop-id 151645

# MLX pack, the documented config (from mlx/, as in Quickstart 4)
python -m fermion_mlx --model ../neutrino-0.6b-chat_v4.bin --mode chat \
    --tokenizer . --rep-penalty 1.05 --gen-tokens 256 --maxseq 1024 \
    --prompt "sup"

Frozen inputs, pinned by sha: judges and banks banks_chatv2.py 0e03e3010f076da28a1058cd1c657caa200cc76e0683f1c17e43c56724e1eb05 (184 facts across 381 phrasings, 100 fragments, 70 threads, 60 beyond-capacity asks, 69 anti-template items); 200 chat hold-outs 482aacb7...; 100 absurd probes 5401f8c6...; identity gold pairs 450a4905.... Preregistration c6a5afe4ef941d08....

Training provenance

Fine-tuned from the predecessor chat state (QAT state sha 211d307f23fd3a1e5c7cb04d9fda7f5ac6f3ae97c4efbe0dfda2e56fb5a00cff) with 300 steps of pure cross-entropy (no KD) at lr 1e-4, batch 8, grad-accum 2, sequence 512, seed 0, on H200 x1. Wall 0.161 h, about $0.87; the whole lane cost about $4 of a $30 cap. Snapshots at 100 / 200 / 300; step 100 shipped.

Corpus: 3,000,610 tokens over eight lanes, every row teacher-generated and gate-verified row by row, exact-deduplicated, rendered serving-exact and EOS-clean. Lane shares as built: fragments 6.7%, common knowledge 11.0%, identity and empathy 4.3%, refusals ~4%, graceful uncertainty 2.2%, small talk 1.9%, anti-template 1.5%, brevity conversations ~68%. Each lane has its own acceptance gate: fragment rows must engage and must not wrap up; fact rows must match a per-fact correct-answer regex; uncertainty rows must carry a humble marker and no fabricated numbers or years; anti-template rows must never emit the identity sentence.

Identity installs cleanly at this scale: 8/8 on every runtime, with manners improved rather than traded.

Weights: get them and verify them

The container ships in this repository. Fetch it and check the digest — MANIFEST.json carries the same sha256, and fermion info verifies it for you automatically:

hf download fermionresearch/Neutrino-0.6B-Chat neutrino-0.6b-chat_v4.bin --local-dir .
shasum -a 256 neutrino-0.6b-chat_v4.bin   # must print 0ec9810f...6272806

The working name for these bytes inside the lane was neutrino-0p6b_v4_chatv2.bin; the two battery receipts in receipts/ still use that name and carry this sha256, so they can be tied to the file you just downloaded.

Limitations

  • General facts are out of scope. The model anchors a curated set of 184 common facts and is honestly uncertain past them, which is what FACTS30 measures. It is not an information source: do not use it for advice, math, code, or tool calling.
  • No knowledge benchmark exists for this checkpoint. Nothing here can be quoted as an MMLU, GSM8K, IFEval, or BFCL result, because none were run. The base card's cells are the nearest legitimate proxy and must be labelled as the base model's.
  • The serving config is load-bearing, per runtime. Serving the raw weights without the documented config reintroduces measured loops: rp 1.00 takes the loop rate from 0.045 to 0.06, and the C binary at plain greedy ignores penalty flags entirely.
  • Greedy decode parity vs the fp reference shows a handful of near-tie argmax flips on the bit-gate prompts, then a greedy cascade; kernel output is fully deterministic across runs and threads (same mechanism as the base SKU's banked analysis).
  • English-centric conversational diet; multilingual behavior is unmeasured here.
  • Refusal training covers absurd, impossible, and mildly sketchy requests. This is a manners layer, NOT a safety system. Do not deploy it as a moderation or safety component.
  • Termination is 0.885 on the C binary and 0.850 on MLX, so roughly one reply in seven reaches the token cap rather than stopping on its own. Set a max_new_tokens you are happy to pay for.

License and attribution

Weights: Apache-2.0. Derivative of Qwen/Qwen3-0.6B (Apache-2.0, Alibaba Cloud), see LICENSE. Chat fine-tune data: ultrachat_200k-derived brevity conversations, a curated common-knowledge bank, synthetic instruction-following, teacher-generated declines of absurd requests, and teacher-gold identity exchanges (receipts in the lane archive).

bin/ binaries: shipped in this repo under bin/, with sha256 sidecars and bin/README.md (EULA: free to use with these weights, no redistribution, no reverse engineering).

Contact: contact@fermionresearch.com

Downloads last month
163
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FermionResearch/Neutrino-0.6B-Chat

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1106)
this model