Tern-0.1

A language model whose weights are born ternary. Every weight matrix in Tern, including the tied embedding and output table, is a code in {−1, 0, +1} from the first step of training to the last. There is no full-precision copy of the weights, no post-training quantization, and no GPU: Tern-0.1 was trained from scratch on a CPU by OkeyMeta Ltd.

The whole model is 959 KB.

Tern-0.1 is the first public checkpoint of the Tern line, released as a proof of the training method, not as a useful assistant. Read the evaluation before you use it.

At a glance

Parameters 3.68M (3.67M ternary, 10K full-precision scales, norms and biases)
Weight format Ternary {−1, 0, +1}, packed 2 bits per weight, one scale per matrix
Activations 8-bit per token
Architecture Looped transformer: 2 unique blocks applied 2 times (depth 4), d_model 256, 4 heads
Positional encoding Compass (see below)
Tokenizer Oji, byte-level BPE, 8,192 tokens + end-of-text
Context 128 tokens
Training data ~12M tokens of FineWeb-Edu
Training hardware 2 CPU threads, no GPU
Optimizer OkeyMeta in-house ternary optimizer (no Adam, no FP latent weights)
File size 959 KB

What is different

Native ternary training. Most 1.58-bit models, BitNet b1.58 included, keep a full-precision "latent" copy of every weight during training and round it to ternary in the forward pass. Tern never stores one. The ternary code itself is the weight, and the optimizer edits codes directly. Training memory for the weights is therefore close to their deployed size.

Looped depth. Two transformer blocks are reused twice, with a learned loop-index vector telling the shared blocks which pass they are on. Depth costs compute, not parameters.

Compass positional encoding. Rotary phases come from three axes instead of one: token order, position inside the current word, and whether the token sits inside a quotation. The model knows where a word starts and where quoted speech begins without having to infer it.

Oji tokenizer. A byte-level BPE trained with a language-parity objective, so text in under-served languages is not split into far more tokens than English.

Evaluation

All models were scored on the same held-out FineWeb-Edu text (92 documents, 302,981 bytes, never seen by Tern) in bits per byte, which is fair across different tokenizers. Lower is better.

Model Params Weight bits bits/byte ↓ LAMBADA (500) ↑
Pythia-160M 162.3M 16 0.978 38.8%
GPT-2 124M 124.4M 16/32 0.991 33.2%
Pythia-70M 70.4M 16 1.107 19.0%
Tern-0.1 3.68M 1.58 2.257 0.0%

Plainly: Tern-0.1 is far behind these baselines. It is 19 to 44 times smaller, stores weights in about a tenth of the bits, and saw roughly 12M training tokens against their hundreds of billions. Its bits-per-byte sits just under a bigram model's (about 2.32 bpb), so it has learned word and local-phrase statistics, not meaning. Its validation loss was still falling when training stopped.

Sample (prompt in bold, temperature 0.8, top-k 40):

The history of science to write they is a find in advived. They also the world say think of members in honths draw things is a gittle slaully in the time the test and the same

English-shaped, not sensible. That is the honest state of a 3.7M-parameter, 12M-token model.

Why release it

Tern-0.1 shows the full pipeline works end to end on commodity hardware: a model trained from random ternary codes to a stable loss curve with no full-precision weights, no GPU and no Adam. The research question for the Tern line is how far that method goes as data and compute grow while weights stay ternary. Tern-0.2, trained on 1.35B tokens, is in progress and will be published against the same benchmark.

Usage

pip install torch safetensors regex huggingface_hub
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("OkeyMetaLtd/Tern-0.1")
sys.path.insert(0, path)
from oji_tokenizer import Oji
from modeling_tern import TernLM

tok = Oji(f"{path}/tokenizer.json")
model = TernLM.from_pretrained(path, tok)
print(model.generate(tok, "Water is made of", max_new_tokens=60, seed=0))

Or from the downloaded folder: python generate.py "Water is made of".

This repository holds inference code only. The forward pass runs ternary weights as float matrices for portability; dedicated ternary kernels would make it much faster.

Limitations

  • Not an assistant. No instruction tuning, no safety tuning, no factual reliability.
  • 128-token context.
  • Trained on English educational web text only, despite a multilingual tokenizer.
  • Will produce ungrammatical and meaningless text.

License

Weights and code are released under CC BY-NC 4.0 for research and non-commercial use. For commercial licensing, contact OkeyMeta Ltd.

Citation

@misc{okeymeta2026tern01,
  title  = {Tern-0.1: A Natively Ternary Language Model Trained on CPU},
  author = {Nwaozor, Okechukwu and {OkeyMeta Ltd}},
  year   = {2026},
  url    = {https://huggingface.co/OkeyMetaLtd/Tern-0.1}
}
Downloads last month
-
Safetensors
Model size
927k params
Tensor type
F32
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train OkeyMetaLtd/Tern-0.1