Nirca Mini
Nirca Mini is a small chat model trained from scratch, with ternary weights:
its linear layers' weights are -1, 0 or +1, with a scale and a bias for each
block of 128 weights. It reads and writes bytes, so it needs no tokenizer. It
is published as it trains: each new checkpoint replaces latest/.
- Small: about 109 million parameters.
- Trained from scratch in ternary, with quantization-aware training from almost the very start; the ternary weights were learned in training, not quantized after.
Nirca was inspired by mini-AGI, and its architecture is deliberately derived from it: Nirca began as a port of mini-AGI to MLX.
➡️ Try it in your browser (WebGPU): https://tg-techie-agents.github.io/nirca-mini-webgpu/mini/
This card is written by program from the checkpoint's own record, at step 69,048 (1,782,983,949 tokens seen).
Model details
- Type: decoder-only language model, chat-tuned, with latent recurrence and a mixture of experts
- Parameters: about 109 million
- Weights: ternary (-1, 0, +1) with a scale and a bias per block of 128 (MLX's 2-bit affine layout); norms and a few small layers in floating point
- Vocabulary: 265 tokens: the 256 byte values and 9 special tokens
- Context length: 2,048 bytes
- Width: 512, with 8 attention heads and rotary position embeddings
- Experts: 32 in one shared pool, 8 active per token
- Recurrence: one block repeated with learned halting, 1 to 24 passes (a mean of 8 in training)
- Chat format: ChatML (
<|im_start|>/<|im_end|>) - Language: English
Architecture
What it keeps from mini-AGI, and what it changes:
- Kept: bytes, no tokenizer; latent recurrence; one shared pool of experts every pass routes into; a halting head; first pretraining text mini-AGI's corpus.
- Changed: ternary weights, trained quantization-aware from step 14,655 of its lineage (its format before that is not recorded); tuned on chats; chat format: ChatML instead of mini-AGI's own tags; a fixed 32-expert pool where mini-AGI grows and prunes its own.
The two dense layers run once. The recurrent block then runs again and again, each pass routing into the same pool of experts, and a halting head decides per byte when another pass would not change the answer (after PonderNet, Banino et al. 2021). Each pass's prediction is weighted by the chance of halting there.
How to use
The demo runs the model in the browser on WebGPU. A page or program loads a checkpoint folder by URL:
latest.json: the newest checkpoint, its step, tokens seen, shape (cfg), and each file's URL, size and sha256.latest/: the newest checkpoint's files.checkpoints/<tokens>k/: every published checkpoint, kept, named by tokens seen in thousands.
Each folder holds model.safetensors (the packed ternary weights) and
config.json (the model's shape and format). Prompts use ChatML:
<|im_start|>user
Hello!<|im_end|>
<|im_start|>assistant
Training data
Across its training, this checkpoint and the runs it was continued from saw:
- An open pretraining corpus of text, in the earlier runs this one was continued from.
- Open chat datasets.
- An open reasoning dataset.
- Synthetic chats from open-weight models, written to give the model its character.
- A small set of private chats, with names, places, contact details and other personal details replaced by placeholders before training.
- Plus 2 datasets not yet public.
Its training data includes text under CC BY 4.0 and CC BY-NC-SA 4.0.
Data and training workflow
Each step is listed because the run's records show it was done:
- ChatML. Every conversation is rendered in ChatML (
<|im_start|>role…<|im_end|>). - Splits by content. Before training, conversations are split into training, validation and test by their content, so one question is never in two splits, and the held-out splits are kept apart from the training machine.
- Length filter. Chat sets are cut to conversations whose whole text fits 2,048 bytes, the model's context.
- Anonymisation. In the private chats, names, places, contact details and other personal details are replaced by typed placeholders, and the result is scanned again before training.
- Identifier check. Synthetic chats that name a known person's identifier are dropped.
- Decontamination. A training conversation that shares an exact turn or an 8-gram with a validation or test set is left out (after PaLM's 8-gram overlap test).
- Loss on the assistant's turns only, from step 61,181: the model learns to answer, not to write the user's turns.
- Quantization-aware training in ternary, from step 14,655 (29,497,375 tokens) on; the steps before it record no weight format. The published weights are the trained ones, not quantized after training.
Training procedure
- Steps: 69,048
- Tokens seen: 1,782,983,949 (one token is one byte)
- Quantization-aware training: from step 14,655 of 69,048, so for 98% of its 1,782,983,949 tokens
- Objective: next-byte prediction; on chats, loss on the assistant's turns only
- Recurrence: the number of passes is drawn per step during training, and the halting head learns when to stop
Evaluation
No benchmark results are published yet. During training, validation loss is measured on held-out conversations that never leave the machine that scores them.
Uses
- Intended: research and experimentation with small, ternary, recurrent models; running a chat model in a browser.
- Out of scope: anything where a wrong answer matters. It is not a source of facts or advice, and it has not been through any safety tuning.
Bias, risks, and limitations
It is a small research model, and it is often wrong or incoherent. It may repeat biases in its training data. It is trained mostly on English, and it has no knowledge of events beyond what its data held. Verify anything it says.
Citation
Nirca Mini's architecture is derived from mini-AGI (see Architecture):
@misc{volotat_miniagi,
author = {volotat},
title = {mini-AGI},
howpublished = {\url{https://github.com/volotat/mini-AGI}},
note = {GitHub repository}
}