File size: 8,884 Bytes
c2ef120
 
49a685e
 
 
 
 
 
 
 
 
 
 
9d480ec
49a685e
 
 
 
 
 
 
 
c2ef120
49a685e
9d480ec
49a685e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9d480ec
49a685e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
9d480ec
 
 
49a685e
 
 
 
 
 
 
 
 
 
 
9d480ec
 
 
49a685e
 
 
 
 
 
 
9d480ec
49a685e
 
 
9d480ec
 
 
49a685e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
---
license: apache-2.0
language: en
tags:
- tinyml
- embedded
- microcontroller
- cortex-m
- risc-v
- xtensa
- rp2040
- esp32
- nrf52840
- avr
- fixed-point
- integer-only
- deterministic
- reproducibility
- quantization
- gpt
- character-level
- custom-code
---

# int-llm `model.mgw` β€” an integer-only GPT that runs bit-identically from x86 servers down to 32-bit microcontrollers, its arithmetic verified down to an 8-bit AVR

This repo hosts the committed reference weights of
**[int-llm](https://github.com/nmicic/int-llm)**: a tiny character-level GPT
trained and sampled **entirely in Q16.48 fixed-point integer arithmetic** β€”
no `float`, no `double`, no `libm` anywhere in the compute path β€” with
**byte-reproducible** results across every tested platform and compiler.

> **Note:** this is a custom C artifact in the project's own `.mgw` container β€”
> **not** a Transformers/safetensors checkpoint. It cannot be loaded with
> `AutoModel` or run on hosted inference; it is consumed by the C programs in
> the GitHub repo below. It is also **not** the TinyLlama-derived inference
> weights discussed in the int-llm write-up β€” this repository publishes only
> the small character-level reference GPT used throughout the project.

Background and motivation: [the int-llm blog post](https://huggingface.co/blog/nmicic/int-llm).

## The file

| | |
|---|---|
| File | `model.mgw` β€” 115,576 bytes |
| sha256 | `466cfe9dba7b888cdaa23dedf4b10351826795793448c8e95dcb0f7a61ed33eb` |
| Model | character-level GPT: `N_EMBD=32, N_HEAD=4, N_LAYER=1`, 14,272 parameters (9 weight tensors) |
| Weights | MGW v1 indexed container; 9 `int64_t` Q16.48 weight tensors plus tokenizer alphabet and RNG state |
| Training | 5000 steps on the public *makemore* names dataset, integer-only (`./gpt_int --save model.mgw`) |
| Math contract | determinism-gate golden hash `c0d933ea340452ec` β€” reproduced by every backend and every validated target below, down to an 8-bit AVR |
| License | Apache-2.0 |

The `.mgw` format (MGW v1) is a small indexed container: a 64-byte header
(magic, version, endianness rejection tag), a 64-byte config block, a tensor
index, the nine weight matrices as raw `int64_t` Q16.48 values, and two
bookkeeping tensors β€” the tokenizer alphabet and the sampler's RNG state
(11 stored tensors in all: 9 learned + 2 bookkeeping).
There is nothing to dequantize: the model was **trained directly in Q16.48**,
so floating-point weights never existed for this model. It can be loaded
from a file (`--load model.mgw`) or **zero-copy from memory**
(`mgpt_load_mem()`), which is how microcontrollers run it straight out of
memory-mapped flash without ever copying the weights into RAM.

## Why host 115 KB of weights?

First, because the size is the point: at 115 KB the *entire model* β€” not a
distilled or re-quantized derivative β€” bakes into the flash of most
mainstream MCUs alongside the code that runs it. Even a 256 KB-flash part
holds both with room to spare (measured firmware sizes below), and on
XIP-capable chips the weights are read in place, occupying zero RAM.

Second, because the file itself is a claim. This *exact byte sequence* is
independently reproduced by running the full 5000-step training on four
different hosts β€” arm64 macOS (clang), x86-64 AMD (gcc), x86-64 Intel (gcc),
and a 2014 Raspberry Pi 1 B+ (32-bit ARMv6, a CPU with no `__int128`) β€” and
every validated inference run on the targets below reproduces the training
host's 20 sampled names **byte-for-byte, PRNG stream included**. Most model
files are "weights we happened to save"; this one is a fixed point (pun
intended) you can re-derive from source.

That makes it useful as an **oracle** for anyone working on quantization,
numerical drift, regression testing, or deterministic inference: any
deviation from it is an implementation bug, never rounding ambiguity β€”
within the documented arithmetic and file-format contract, there is no
floating-point tolerance to hide behind.

## Validated on real hardware

The train-big/run-small loop closes on microcontrollers: train on a laptop,
run inference-only from this weight file on a $5 board. Every target below
passed both checks (determinism-grid golden hash + byte-identical 20-sample
inference), except the 8-bit Mega 2560 β€” a board over a decade old β€” which
runs the determinism gate only (8 KB RAM), computing the same 64-bit hash
from 8-bit ALU instructions; raw provenance-stamped transcripts live in
[`validation/cpu/`](https://github.com/nmicic/int-llm/tree/main/validation/cpu).

| target | ISA | determinism | 20 samples |
|---|---|---|---|
| XIAO RP2040 (Cortex-M0+ @ 133 MHz) | Armv6-M | 10.9 s | 6.7 s |
| Raspberry Pi Pico 2 (RP2350, ARM mode, Cortex-M33) | Armv8-M | 4.1 s | 3.1 s |
| Raspberry Pi Pico 2 (RP2350, RISC-V mode, Hazard3) | rv32imac | 5.4 s | 3.8 s |
| ESP32-C6 | rv32imac | 5.6 s | 2.0 s |
| Heltec V3 (ESP32-S3, LX7 @ 240 MHz) | Xtensa | 3.9 s | 1.1 s |
| LILYGO T-Beam (ESP32, LX6 @ 240 MHz) | Xtensa | 4.2 s | 2.6 s |
| XIAO nRF52840 (Cortex-M4F @ 64 MHz) | Armv7E-M | 13.9 s | 3.4 s |
| Arduino MKR Zero (SAMD21, Cortex-M0+ @ 48 MHz, 256 KB flash / 32 KB RAM) | Armv6-M | 37.2 s | 26.4 s |
| Arduino Mega 2560 (ATmega2560, 8-bit AVR @ 16 MHz, 8 KB RAM) | AVR | 747.9 s | n/a (8 KB RAM) |
| Raspberry Pi 1 B+ (2014, ARMv6, 32-bit Linux) | ARMv6 | 0.7 s | 0.2 s |
| AMD Ryzen 7 7700 (Linux, gcc) | x86-64 | native + portable | byte-identical |
| Intel Core i7-7700 (Linux, gcc) | x86-64 | native + portable | byte-identical |

Same die, two ISAs: the Pico 2 reproduces the identical output in both its
ARM and RISC-V boot modes. The Linux rows additionally rerun the **full
training** and reproduce this repo's `model.mgw` byte-for-byte.

**Footprint** (XIAO RP2040): the complete firmware is
187,932 B of flash β€” *including* the whole 115 KB weight file baked into
rodata β€” and 15,268 B of static RAM (5.8% of the RP2040's 264 KB). Weights
are read in place over XIP flash; RAM holds only KV cache, activations, and
the USB stack. The 256 KB claim is measured, not extrapolated: on the
Arduino MKR Zero's SAMD21 (256 KB flash / 32 KB RAM), code plus the whole
weight file build to 133,472 B β€” 50.9% of flash β€” with 13,124 B of RAM.

## Run it

```bash
git clone https://github.com/nmicic/int-llm
cd int-llm
make gpt_int
./gpt_int --load model.mgw     # 20 names, byte-identical on every validated target
```

The identical file is committed in the GitHub repo; to fetch just the
weights from here instead:

```bash
hf download nmicic/int-llm model.mgw --local-dir .
# for strict reproducibility, pin a revision once published:
#   hf download nmicic/int-llm model.mgw --revision <commit-or-tag> --local-dir .
```

To retrain and verify the reproducibility claim yourself:

```bash
make input                      # fetch the names dataset
./gpt_int --save model2.mgw     # ~2 s on a desktop, ~10 min on a Pi 1
cmp model2.mgw model.mgw        # exit 0
```

The MCU harnesses (PlatformIO, one folder per board, flash + serial-capture
scripts) are in
[`validation/cpu/`](https://github.com/nmicic/int-llm/tree/main/validation/cpu).

## Intended use

This artifact is intended for: validating integer-only or fixed-point
implementations against an exact reference, regression testing, deterministic
inference research, and embedded/TinyML experiments. It is **not** intended
for production language generation, hosted inference, use with the
Transformers library, or capability comparisons against modern LLMs.

## Limitations & provenance

- **"Integer-only" refers to the model compute path**: training, inference,
  sampling, and serialization are all Q16.48 integer arithmetic. Build
  tooling, logging, and timing on the host are outside that claim.
- This is a **toy by design**: a character-level *name generator*, not a
  general text LLM. `BLOCK_SIZE` is 8, so generated names are at most
  8 characters. Its value is the methodology β€” exact integer arithmetic and
  cross-platform byte-reproducibility β€” not the language modeling.
- Training data: [`names.txt` from Andrej Karpathy's makemore](https://github.com/karpathy/makemore)
  (MIT license), fetched by `make input`;
  sha256 `0a30b5557f192f32ab962680889aac5f6fda0f4cecf40a6d0b5694f58ea8cc4d`.
- MGW v1 is a **host-native-endian** format with an endianness rejection tag
  in the header; this published file is **little-endian** (as are all
  validated targets).
- The loader is a research-grade parser, not hardened against adversarial
  inputs β€” verify the sha256 above and load only trusted `.mgw` files.

## Links

- Source, math library, and validation records: **https://github.com/nmicic/int-llm**
- Write-up: **https://huggingface.co/blog/nmicic/int-llm**
- The portable Q16.48 dual-backend math header matured in the sibling project **[astro-nav-int](https://github.com/nmicic/astro-nav-int)**