stories15M-ja
A 15M-parameter Llama that writes short Japanese children's stories: exactly the shape of
karpathy/tinyllamas' stories15M.bin
(dim=288, n_layers=6, n_heads=6, hidden_dim=768, vocab_size=32000, max_seq_len=256, tied
classifier), trained from scratch on
shibatch/TinyStories-JA with a Japanese
vocabulary of its own. Made for transformer-dojo, whose
participants implement this model's forward pass themselves: swapping in these weights and this
tokenizer makes their finished code speak Japanese.
ๆฏๆฅใใใใฏในใฏใๆฏใใใจไธ็ทใซใใๅบญใฎๅคใง้ใใงใใพใใใใใใฏในใฏใ่ตฐใๅใฃใใใใธใฃใณใใใใใๆขๆคใใใใใใฎใๅคงๅฅฝใใงใใใ ใใๆฅใฎใใจใใใใฏในใฏใๅบญใฎไธญใซใๅคงใใใฆใใใใใฎใ่ฆใคใใพใใใใใใฏใๅคงใใชใใใงใใ๏ผใใใฏในใฏๆใใชใฃใฆใๆณฃใๅบใใฆใใพใใพใใใ ใๆฏใใใฏใใใฏในใใใ ใฃใจๆฑใใใใฆใใๅคงไธๅคซใใใใใฏในใใใใฏใใ ใฎใใใใใใชใใๅทใคใใใใใชใใใใจ่จใใพใใใ ใใใฏในใฏใพใ ๆใใฃใใใใฉใใๆฏใใใฎใใจใไฟกใใฆใใพใใใใๆฏใใใฏใใใฎใใใฏใใ ใๅ้ใๆขใใฆใใใ ใใชใฎใ ใจ่จใใพใใใ ใใใฏในใฏๅฎๅฟใใฆใใพใใๅบญใง้ใณๅงใใพใใใใใใฏในใฏใใใฎใใใใใใจใๅ้ใซใชใใใใไบๅบฆใจๆใใใใจใฏใใใพใใใงใใใ
(greedy, INT4, prompt ใๆฏๆฅใใใใฏในใฏใ)
Files
| File | What it is | sha256 |
|---|---|---|
model.bin |
fp32 weights, llama2.c's v0 checkpoint format (what run.c reads) |
1b4b477270501e9f47dcf24ba65777546422dbbb663ee4fdaa77d8f7c9e90438 |
tokenizer.bin |
the vocabulary in llama2.c's tokenizer.bin format |
f5d3fa64ca1fc90eebbe998ba9254ebe93f1dfb3aa8d96d0ced657dbb9912a15 |
tokenizer.model |
the same vocabulary as a sentencepiece model | c6cb86ade40e14c1cb691bf32105cae52f1fdd6100dce815e62d57ceee39d3bf |
train_config.json, log.jsonl |
the training run's configuration and its evaluation log |
model.bin and tokenizer.bin run unchanged in llama2.c:
./run model.bin -z tokenizer.bin -i "ใใใใใใใ". transformer-dojo quantizes model.bin to
its INT4 format itself; the INT4 file is not published here.
Training
- Data:
shibatch/TinyStories-JAat revision155b4e5842076c90e07a4934d5b68410fe21f65f, fieldtext_ja, train split only; records that are empty, shorter than 20 characters, without hiragana, over 5% Latin letters, or containing Hangul, Thai, Arabic or Cyrillic are dropped (0.04%). 329M tokens per epoch. - Vocabulary: sentencepiece BPE, 32000 pieces, byte fallback, trained on the first 200,000
stories; llama2.c's layout (
<unk>, BOS, EOS, 256 byte pieces, then merges), withremove_extra_whitespacesoff so it tokenizes as llama2.c's encoder does. - Run: 8 epochs (2.64B tokens), peak learning rate 2e-3. AdamW (ฮฒโ 0.95, weight decay 0.1), linear warmup 1000 steps then cosine decay to 0, batch 128 ร 4 ร 256 tokens, bf16 autocast, one RTX 5090 for 48 minutes.
- Result: validation loss 2.305. Quantized to transformer-dojo's INT4 (group size 32), held-out NLL on 40 validation stories rises 2.034 โ 2.160 (+6.2%), and no greedy continuation of 20 prompts repeats a 4-gram four times running.
The code is in transformer-dojo's tools/train/ (trimmed from llama2.c); the full log of runs is
its docs/plan/m-ja-training-log.md.
Limits
A model this small forgets its own plot: names drift, and a story can circle back on a sentence. The Japanese is the dataset's โ machine-translated English stories, so the register is translationese in places (ๅฝผ/ๅฝผๅฅณ, ใใซใฃใใ็ฌใฃใฆ่จใใพใใใ). It knows children's-story vocabulary only; other text is tokenized fine (byte fallback) but not understood.
Licence and provenance
- Weights and tokenizer: MIT.
- Data: TinyStories (Ronen Eldan, Yuanzhi Li; arXiv:2305.07759), translated to Japanese by Naoki Shibata, CDLA-Sharing-1.0. A trained model is a "Result" of the data (ยง1.11), and ยง3.5: "This Agreement imposes no obligations or restrictions on Your Use or Publication of Results."
- Translation model: the translations were made primarily with
google/gemma-4-26B-A4B, Apache-2.0, which places no obligation on outputs. - Code: llama2.c by Andrej Karpathy, MIT.