File size: 3,790 Bytes
96a11f8
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
---
license: mit
tags:
- bitnet
- ternary
- from-scratch
- tinystories
- pretraining
---

# RivaQuant

A small (162M-param), from-scratch decoder-only transformer with **BitNet
b1.58 ternary weights** ({-1, 0, 1}) in every attention/MLP projection,
trained on [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories).
Architecture and training code: [entropy-om/rivaquant](https://github.com/entropy-om/rivaquant).

RivaQuant is the project/repo name, not a self-chosen identity — see
"Does it know its own name?" below for why.

## What this actually is

- **Architecture**: BitNet b1.58 ternary linear layers (weight quantization
  via absmean + sign, per-token int8 activation quantization, straight-through
  estimator for gradients), RoPE, standard decoder-only transformer block
  structure otherwise. Adapted from real reference implementations
  (kyegomez/BitNet's math, Microsoft's BitNet b1.58 paper), not reinvented
  from memory.
- **Params**: 162,213,888
- **Trained**: 17,500 steps (best checkpoint by
  validation loss, out of 20,000 total steps run) on TinyStories, batch
  size 8 with 4x gradient accumulation (effective batch 32), block size 256.
- **Validation perplexity**: 5.455

## What this is not

Not a general-purpose assistant, not instruction-tuned, not evaluated on
anything beyond TinyStories-style short story completion. Read the actual
architecture and eval code before drawing conclusions from perplexity alone.

## Does it know its own name?

No. Prompted with "My name is" / "I am called" / "You can call me" (15
samples, unfiltered), it produces a different plausible children's-story
character name every time — a direct artifact of training purely on
TinyStories, which is full of characters introducing themselves that way.
There is no consistent self-identity to report, so none is claimed:

- `My name is Max. What is your name?"

Lily said, "That is Max, the owl`
- `My name is Daisy. I like your hat."

Mia smiled and said, "That is very nice`
- `My name is Lily. I live in a small village. Do you want to play with me?"

L`
- `My name is Lily," said Lily.

The boy said, "I'm happy to meet you. I`
- `My name is Ben. What is your name?"

Ben says, "I am Ben, and this is`
- `I am called it. I love you. You are my best friend."

She hugs Ben and says,`
- `I am called Tim. I like to play with you." They played and had fun together. Lily learned that sharing`
- `I am called a horse. I am looking for a home. Do you want to come with me?"

`
- `I am called Tom. I want to play with you. But you have to wait for me. I am taking`
- `I am called an elephant. I like to eat honey. Do you want a banana?"

Lily and`
- `You can call me if you want. I will give you a hug."

Lily smiled and said, "`
- `You can call me. I have a phone. It is called my phone. Do you like it?"

Ben`
- `You can call me and get up, but please stay down," said the dog.

Mama smiled and said`
- `You can call me and ask me if you want to come. There's no need to do you."

Anna`
- `You can call me, or I will call your dad."

Lily and Ben went back to the fence.`

## Training bugs hit and fixed along the way

1. `torch.nn.RMSNorm` requires torch>=2.4; the training environment shipped
   2.1.0. Custom RMSNorm module, same math, no version dependency.
2. CUDA OOM at batch=32/block=512 on a 24GB card — BitNet's straight-through
   estimator keeps extra activation copies per layer, more memory-hungry
   than plain `nn.Linear` at the same nominal size. Fixed with a smaller
   micro-batch + gradient accumulation to keep the same effective batch.

Both caught mid-training by an automated cost-safety watcher that stops the
GPU on any crash signature, understood, fixed, verified with a smoke test,
and relaunched — not discovered after the fact.