File size: 3,790 Bytes
96a11f8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 | ---
license: mit
tags:
- bitnet
- ternary
- from-scratch
- tinystories
- pretraining
---
# RivaQuant
A small (162M-param), from-scratch decoder-only transformer with **BitNet
b1.58 ternary weights** ({-1, 0, 1}) in every attention/MLP projection,
trained on [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories).
Architecture and training code: [entropy-om/rivaquant](https://github.com/entropy-om/rivaquant).
RivaQuant is the project/repo name, not a self-chosen identity — see
"Does it know its own name?" below for why.
## What this actually is
- **Architecture**: BitNet b1.58 ternary linear layers (weight quantization
via absmean + sign, per-token int8 activation quantization, straight-through
estimator for gradients), RoPE, standard decoder-only transformer block
structure otherwise. Adapted from real reference implementations
(kyegomez/BitNet's math, Microsoft's BitNet b1.58 paper), not reinvented
from memory.
- **Params**: 162,213,888
- **Trained**: 17,500 steps (best checkpoint by
validation loss, out of 20,000 total steps run) on TinyStories, batch
size 8 with 4x gradient accumulation (effective batch 32), block size 256.
- **Validation perplexity**: 5.455
## What this is not
Not a general-purpose assistant, not instruction-tuned, not evaluated on
anything beyond TinyStories-style short story completion. Read the actual
architecture and eval code before drawing conclusions from perplexity alone.
## Does it know its own name?
No. Prompted with "My name is" / "I am called" / "You can call me" (15
samples, unfiltered), it produces a different plausible children's-story
character name every time — a direct artifact of training purely on
TinyStories, which is full of characters introducing themselves that way.
There is no consistent self-identity to report, so none is claimed:
- `My name is Max. What is your name?"
Lily said, "That is Max, the owl`
- `My name is Daisy. I like your hat."
Mia smiled and said, "That is very nice`
- `My name is Lily. I live in a small village. Do you want to play with me?"
L`
- `My name is Lily," said Lily.
The boy said, "I'm happy to meet you. I`
- `My name is Ben. What is your name?"
Ben says, "I am Ben, and this is`
- `I am called it. I love you. You are my best friend."
She hugs Ben and says,`
- `I am called Tim. I like to play with you." They played and had fun together. Lily learned that sharing`
- `I am called a horse. I am looking for a home. Do you want to come with me?"
`
- `I am called Tom. I want to play with you. But you have to wait for me. I am taking`
- `I am called an elephant. I like to eat honey. Do you want a banana?"
Lily and`
- `You can call me if you want. I will give you a hug."
Lily smiled and said, "`
- `You can call me. I have a phone. It is called my phone. Do you like it?"
Ben`
- `You can call me and get up, but please stay down," said the dog.
Mama smiled and said`
- `You can call me and ask me if you want to come. There's no need to do you."
Anna`
- `You can call me, or I will call your dad."
Lily and Ben went back to the fence.`
## Training bugs hit and fixed along the way
1. `torch.nn.RMSNorm` requires torch>=2.4; the training environment shipped
2.1.0. Custom RMSNorm module, same math, no version dependency.
2. CUDA OOM at batch=32/block=512 on a 24GB card — BitNet's straight-through
estimator keeps extra activation copies per layer, more memory-hungry
than plain `nn.Linear` at the same nominal size. Fixed with a smaller
micro-batch + gradient accumulation to keep the same effective batch.
Both caught mid-training by an automated cost-safety watcher that stops the
GPU on any crash signature, understood, fixed, verified with a smoke test,
and relaunched — not discovered after the fact.
|