File size: 1,215 Bytes
e5ab840
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
---
language: en
pipeline_tag: text-generation
tags:
- assembly
- from-scratch
- chat
---

# Mnemonic 126M

A small chat model where every part (data pipeline, tokenizer, GPU training, quantization and inference) is hand-written assembly: x86-64 on the CPU, PTX on the GPU, WebAssembly in the browser.

Code: https://github.com/Supergoatscriptguy/Mnemonic

| file | weights | size |
|---|---|---|
| `mnemonic-q8.mnm` | int8, one scale per row | 121 MB |
| `mnemonic-q4.mnm` | int4, one scale per group of 32 | 75 MB |

**Model.** Llama-style decoder: 16 layers, d_model 768, 12 query / 4 key-value heads, SwiGLU (ffn 2048), RoPE, RMSNorm, tied embeddings, no biases. Context 1024, byte-level BPE vocab of 32768.

**Training.** Pretrained on 5B tokens of FineWeb-Edu on one RTX 5070 Ti (19 hours), then fine-tuned for chat on smol-smoltalk plus a small identity set.

**Format.** `.mnm` is the project's own format (a 4 KB header, then the tensors), read by the engines in the repo (`chat/engine.asm`, `site/engine.wat`). It is not a transformers checkpoint.

**Limitations.** It's a small model: it makes confident mistakes, is weak at math, and loses the thread in long conversations.

Made by Supergoatscriptguy.