algo-v1-codepoint
A 38.8M-parameter decoder-only transformer, part-way through training.
This checkpoint is undertrained and not a usable assistant. It is published as a research artifact: the code, the training recipe, and an honest account of what 16.4M tokens bought. See Status before using it.
Architecture
| Parameters | 38,806,016 (tied embeddings) |
| Layers | 8 |
d_model |
512 |
| Attention | Grouped-query attention, 8 query heads / 2 KV heads (4:1), head_dim 64 |
| FFN | SwiGLU, ffn_dim 2048 |
| Context length | 8,192 |
| Precision | fp32 |
RoPE positional encoding, RMSNorm, pre-norm blocks. The tied output projection is 82.7% of the per-layer parameter count.
Tokenizer
Code-point level (Tokenizer in algo/data.py), not byte level:
id = ord(char) + 6, with 6 reserved ids (pad=0, eos=1, unk=2, tool=3,
python=4, shell=5). vocab_size = 16384; characters beyond that map to
unk_id. Multi-byte characters therefore cost more than one token in practice
only when they exceed the ceiling โ most map to a single id.
python -m algo.train and algo.data.Tokenizer(vocab_size=16384) are the
reference implementations. The class is a frozen dataclass with no learned
vocabulary file, so nothing needs to be downloaded to reconstruct it.
A later revision of this project moved to a byte-level tokenizer with
vocab_size=262. Those weights are not interchangeable with this one: id 200 means a different character here than there. Use thealgo-v1-codepointgit branch, which pins the tokenizer this checkpoint was trained with.
Training
Corpus: 50,000,000 tokens built from Hugging Face (Indonesian Wikipedia 23.5M /
English Wikipedia 23.5M / GSM8K 3.0M), content digests pinned in
data/manifest.json and enforced on every build.
| Tokens seen | 16,384,000 (32.8% of one epoch) |
| Steps | 2,000 of 6,104 in the cosine schedule |
| Hardware | 2x Tesla T4 (DataParallel) |
| Throughput | 10,062 tokens/sec |
| Wall clock | ~29 minutes |
| Final training loss | 2.1023 |
| Optimizer | AdamW, lr 2e-3 cosine, warmup 500, weight decay 0.1, grad clip 1.0 |
| Batch | 8,192 tokens (8 sequences x 1,024) |
The run stopped at 32.8% of the schedule, so the learning rate never annealed โ it was still near 0.0017 at the final step against a 0.002 peak. This checkpoint is a mid-schedule slice, not a converged model.
Status
Loss and perplexity measured by scripts/baseline.py over 8 windows per split:
| Split | Loss | Perplexity |
|---|---|---|
| Seen | 2.066 | 7.895 |
| Held-out | 2.181 | 8.855 |
Train/held-out gap 0.115 +/- 0.085 (t = 1.35 on 8 degrees of freedom) โ not statistically distinguishable from zero, so no evidence of memorisation at this scale. Equally, no evidence of generalisation either: 16.4M tokens is a small fraction of an epoch.
Greedy samples from the checkpoint, showing the model has learned surface form without learning to continue coherently:
"Ibukota Indonesia adalah" -> " di peringkan di perina peringkan di peringkan p"
"def fibonacci(n):" -> "a seced an the secord the spersity of the spersi"
Language identity is preserved (Indonesian in, Indonesian out), and GSM8K
formatting artefacts such as the <<...>> calculator delimiters appear. Longer
coherence, facts, and arithmetic do not.
Intended use
Research and education: studying training recipes, tokenizer choices, and generalisation behaviour at small scale. Not suitable as an assistant, and not evaluated for any safety property.
Reproducing
The full pipeline โ corpus construction, training, evaluation โ is in the
repository under the algo-v1-codepoint branch. See docs/OPERATIONS.md.
Training was performed on Kaggle; the corpus is rebuilt from the Hub on each run rather than committed, since it is 100 MB of derived data.
License
ISC. See LICENSE.
Corpora keep their own upstream licenses and are not redistributed here: Wikipedia is CC BY-SA 4.0, GSM8K is MIT.