algo-v1-codepoint

A 38.8M-parameter decoder-only transformer, part-way through training.

This checkpoint is undertrained and not a usable assistant. It is published as a research artifact: the code, the training recipe, and an honest account of what 16.4M tokens bought. See Status before using it.

Architecture

Parameters 38,806,016 (tied embeddings)
Layers 8
d_model 512
Attention Grouped-query attention, 8 query heads / 2 KV heads (4:1), head_dim 64
FFN SwiGLU, ffn_dim 2048
Context length 8,192
Precision fp32

RoPE positional encoding, RMSNorm, pre-norm blocks. The tied output projection is 82.7% of the per-layer parameter count.

Tokenizer

Code-point level (Tokenizer in algo/data.py), not byte level: id = ord(char) + 6, with 6 reserved ids (pad=0, eos=1, unk=2, tool=3, python=4, shell=5). vocab_size = 16384; characters beyond that map to unk_id. Multi-byte characters therefore cost more than one token in practice only when they exceed the ceiling โ€” most map to a single id.

python -m algo.train and algo.data.Tokenizer(vocab_size=16384) are the reference implementations. The class is a frozen dataclass with no learned vocabulary file, so nothing needs to be downloaded to reconstruct it.

A later revision of this project moved to a byte-level tokenizer with vocab_size=262. Those weights are not interchangeable with this one: id 200 means a different character here than there. Use the algo-v1-codepoint git branch, which pins the tokenizer this checkpoint was trained with.

Training

Corpus: 50,000,000 tokens built from Hugging Face (Indonesian Wikipedia 23.5M / English Wikipedia 23.5M / GSM8K 3.0M), content digests pinned in data/manifest.json and enforced on every build.

Tokens seen 16,384,000 (32.8% of one epoch)
Steps 2,000 of 6,104 in the cosine schedule
Hardware 2x Tesla T4 (DataParallel)
Throughput 10,062 tokens/sec
Wall clock ~29 minutes
Final training loss 2.1023
Optimizer AdamW, lr 2e-3 cosine, warmup 500, weight decay 0.1, grad clip 1.0
Batch 8,192 tokens (8 sequences x 1,024)

The run stopped at 32.8% of the schedule, so the learning rate never annealed โ€” it was still near 0.0017 at the final step against a 0.002 peak. This checkpoint is a mid-schedule slice, not a converged model.

Status

Loss and perplexity measured by scripts/baseline.py over 8 windows per split:

Split Loss Perplexity
Seen 2.066 7.895
Held-out 2.181 8.855

Train/held-out gap 0.115 +/- 0.085 (t = 1.35 on 8 degrees of freedom) โ€” not statistically distinguishable from zero, so no evidence of memorisation at this scale. Equally, no evidence of generalisation either: 16.4M tokens is a small fraction of an epoch.

Greedy samples from the checkpoint, showing the model has learned surface form without learning to continue coherently:

"Ibukota Indonesia adalah"  -> " di peringkan di perina peringkan di peringkan p"
"def fibonacci(n):"         -> "a seced an the secord the spersity of the spersi"

Language identity is preserved (Indonesian in, Indonesian out), and GSM8K formatting artefacts such as the <<...>> calculator delimiters appear. Longer coherence, facts, and arithmetic do not.

Intended use

Research and education: studying training recipes, tokenizer choices, and generalisation behaviour at small scale. Not suitable as an assistant, and not evaluated for any safety property.

Reproducing

The full pipeline โ€” corpus construction, training, evaluation โ€” is in the repository under the algo-v1-codepoint branch. See docs/OPERATIONS.md.

Training was performed on Kaggle; the corpus is rebuilt from the Hub on each run rather than committed, since it is 100 MB of derived data.

License

ISC. See LICENSE.

Corpora keep their own upstream licenses and are not redistributed here: Wikipedia is CC BY-SA 4.0, GSM8K is MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support