Text Generation
Japanese
llama2.c
tinystories
japanese

stories15M-ja

A 15M-parameter Llama that writes short Japanese children's stories: exactly the shape of karpathy/tinyllamas' stories15M.bin (dim=288, n_layers=6, n_heads=6, hidden_dim=768, vocab_size=32000, max_seq_len=256, tied classifier), trained from scratch on shibatch/TinyStories-JA with a Japanese vocabulary of its own. Made for transformer-dojo, whose participants implement this model's forward pass themselves: swapping in these weights and this tokenizer makes their finished code speak Japanese.

ๆฏŽๆ—ฅใ€ใƒžใƒƒใ‚ฏใ‚นใฏใŠๆฏใ•ใ‚“ใจไธ€็ท’ใซใ€ใŠๅบญใฎๅค–ใง้Šใ‚“ใงใ„ใพใ—ใŸใ€‚ใƒžใƒƒใ‚ฏใ‚นใฏใ€่ตฐใ‚Šๅ›žใฃใŸใ‚Šใ€ใ‚ธใƒฃใƒณใƒ—ใ—ใŸใ‚Šใ€ๆŽขๆคœใ—ใŸใ‚Šใ™ใ‚‹ใฎใŒๅคงๅฅฝใใงใ—ใŸใ€‚ ใ‚ใ‚‹ๆ—ฅใฎใ“ใจใ€ใƒžใƒƒใ‚ฏใ‚นใฏใŠๅบญใฎไธญใซใ€ๅคงใใใฆใ“ใ‚ใ„ใ‚‚ใฎใ‚’่ฆ‹ใคใ‘ใพใ—ใŸใ€‚ใใ‚Œใฏใ€ๅคงใใชใƒ˜ใƒ“ใงใ—ใŸ๏ผใƒžใƒƒใ‚ฏใ‚นใฏๆ€–ใใชใฃใฆใ€ๆณฃใๅ‡บใ—ใฆใ—ใพใ„ใพใ—ใŸใ€‚ ใŠๆฏใ•ใ‚“ใฏใƒžใƒƒใ‚ฏใ‚นใ‚’ใŽใ‚…ใฃใจๆŠฑใใ—ใ‚ใฆใ€ใ€Œๅคงไธˆๅคซใ‚ˆใ€ใƒžใƒƒใ‚ฏใ‚นใ€‚ใ‚ใ‚ŒใฏใŸใ ใฎใƒ˜ใƒ“ใ‚ˆใ€‚ใ‚ใชใŸใ‚’ๅ‚ทใคใ‘ใŸใ‚Šใ—ใชใ„ใ‚ใ€ใจ่จ€ใ„ใพใ—ใŸใ€‚ ใƒžใƒƒใ‚ฏใ‚นใฏใพใ ๆ€–ใ‹ใฃใŸใ‘ใ‚Œใฉใ€ใŠๆฏใ•ใ‚“ใฎใ“ใจใ‚’ไฟกใ˜ใฆใ„ใพใ—ใŸใ€‚ใŠๆฏใ•ใ‚“ใฏใ€ใใฎใƒ˜ใƒ“ใฏใŸใ ใŠๅ‹้”ใ‚’ๆŽขใ—ใฆใ„ใ‚‹ใ ใ‘ใชใฎใ ใจ่จ€ใ„ใพใ—ใŸใ€‚ ใƒžใƒƒใ‚ฏใ‚นใฏๅฎ‰ๅฟƒใ—ใฆใ€ใพใŸใŠๅบญใง้Šใณๅง‹ใ‚ใพใ—ใŸใ€‚ใƒžใƒƒใ‚ฏใ‚นใฏใ€ใใฎใ“ใ‚ใ„ใƒ˜ใƒ“ใจใŠๅ‹้”ใซใชใ‚Šใ€ใ‚‚ใ†ไบŒๅบฆใจๆ€–ใŒใ‚‹ใ“ใจใฏใ‚ใ‚Šใพใ›ใ‚“ใงใ—ใŸใ€‚

(greedy, INT4, prompt ใ€ŒๆฏŽๆ—ฅใ€ใƒžใƒƒใ‚ฏใ‚นใฏใ€)

Files

File What it is sha256
model.bin fp32 weights, llama2.c's v0 checkpoint format (what run.c reads) 1b4b477270501e9f47dcf24ba65777546422dbbb663ee4fdaa77d8f7c9e90438
tokenizer.bin the vocabulary in llama2.c's tokenizer.bin format f5d3fa64ca1fc90eebbe998ba9254ebe93f1dfb3aa8d96d0ced657dbb9912a15
tokenizer.model the same vocabulary as a sentencepiece model c6cb86ade40e14c1cb691bf32105cae52f1fdd6100dce815e62d57ceee39d3bf
train_config.json, log.jsonl the training run's configuration and its evaluation log

model.bin and tokenizer.bin run unchanged in llama2.c: ./run model.bin -z tokenizer.bin -i "ใ‚€ใ‹ใ—ใ‚€ใ‹ใ—ใ€". transformer-dojo quantizes model.bin to its INT4 format itself; the INT4 file is not published here.

Training

  • Data: shibatch/TinyStories-JA at revision 155b4e5842076c90e07a4934d5b68410fe21f65f, field text_ja, train split only; records that are empty, shorter than 20 characters, without hiragana, over 5% Latin letters, or containing Hangul, Thai, Arabic or Cyrillic are dropped (0.04%). 329M tokens per epoch.
  • Vocabulary: sentencepiece BPE, 32000 pieces, byte fallback, trained on the first 200,000 stories; llama2.c's layout (<unk>, BOS, EOS, 256 byte pieces, then merges), with remove_extra_whitespaces off so it tokenizes as llama2.c's encoder does.
  • Run: 8 epochs (2.64B tokens), peak learning rate 2e-3. AdamW (ฮฒโ‚‚ 0.95, weight decay 0.1), linear warmup 1000 steps then cosine decay to 0, batch 128 ร— 4 ร— 256 tokens, bf16 autocast, one RTX 5090 for 48 minutes.
  • Result: validation loss 2.305. Quantized to transformer-dojo's INT4 (group size 32), held-out NLL on 40 validation stories rises 2.034 โ†’ 2.160 (+6.2%), and no greedy continuation of 20 prompts repeats a 4-gram four times running.

The code is in transformer-dojo's tools/train/ (trimmed from llama2.c); the full log of runs is its docs/plan/m-ja-training-log.md.

Limits

A model this small forgets its own plot: names drift, and a story can circle back on a sentence. The Japanese is the dataset's โ€” machine-translated English stories, so the register is translationese in places (ๅฝผ/ๅฝผๅฅณ, ใ€Œใซใฃใ“ใ‚Š็ฌ‘ใฃใฆ่จ€ใ„ใพใ—ใŸใ€). It knows children's-story vocabulary only; other text is tokenized fine (byte fallback) but not understood.

Licence and provenance

  • Weights and tokenizer: MIT.
  • Data: TinyStories (Ronen Eldan, Yuanzhi Li; arXiv:2305.07759), translated to Japanese by Naoki Shibata, CDLA-Sharing-1.0. A trained model is a "Result" of the data (ยง1.11), and ยง3.5: "This Agreement imposes no obligations or restrictions on Your Use or Publication of Results."
  • Translation model: the translations were made primarily with google/gemma-4-26B-A4B, Apache-2.0, which places no obligation on outputs.
  • Code: llama2.c by Andrej Karpathy, MIT.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for opsbr/stories15M-ja

Finetunes
1 model

Dataset used to train opsbr/stories15M-ja

Paper for opsbr/stories15M-ja