|
Download README.md from SuperGoatScriptGuy/Mnemonic: direct link, hf CLI and curl.
- Browser
- Download file 1.22 kB
-
https://huggingface.co/SuperGoatScriptGuy/Mnemonic/resolve/main/README.md
- Command line
-
hf download hf://SuperGoatScriptGuy/Mnemonic/README.md
-
curl -L -o README.md https://huggingface.co/SuperGoatScriptGuy/Mnemonic/resolve/main/README.md
1.22 kB
| language: en | |
| pipeline_tag: text-generation | |
| tags: | |
| - assembly | |
| - from-scratch | |
| - chat | |
| # Mnemonic 126M | |
| A small chat model where every part (data pipeline, tokenizer, GPU training, quantization and inference) is hand-written assembly: x86-64 on the CPU, PTX on the GPU, WebAssembly in the browser. | |
| Code: https://github.com/Supergoatscriptguy/Mnemonic | |
| | file | weights | size | | |
| |---|---|---| | |
| | `mnemonic-q8.mnm` | int8, one scale per row | 121 MB | | |
| | `mnemonic-q4.mnm` | int4, one scale per group of 32 | 75 MB | | |
| **Model.** Llama-style decoder: 16 layers, d_model 768, 12 query / 4 key-value heads, SwiGLU (ffn 2048), RoPE, RMSNorm, tied embeddings, no biases. Context 1024, byte-level BPE vocab of 32768. | |
| **Training.** Pretrained on 5B tokens of FineWeb-Edu on one RTX 5070 Ti (19 hours), then fine-tuned for chat on smol-smoltalk plus a small identity set. | |
| **Format.** `.mnm` is the project's own format (a 4 KB header, then the tensors), read by the engines in the repo (`chat/engine.asm`, `site/engine.wat`). It is not a transformers checkpoint. | |
| **Limitations.** It's a small model: it makes confident mistakes, is weak at math, and loses the thread in long conversations. | |
| Made by Supergoatscriptguy. | |