Kanjiland — Translation model (Ja→En)

A 52.3M encoder-decoder Transformer built from scratch in raw PyTorch (no HuggingFace transformers in the modeling path). Part of Kanjiland.

  • Architecture: 6+6 layers, d_model 512, 8 heads, RoPE, pre-LN, three-way tied embeddings. Ablation-confirmed config (RoPE / 16k / tied).
  • Training: 100k steps on a 22.1M-pair Ja-En corpus, single RTX 4090.
  • Eval (beam 4, chrF/BLEU/COMET): KFTT-test 47.2 / 20.5 / 0.77; m2-test 55.3 / 30.4 / 0.81; WMT22 42.2 / 16.3 / 0.76 (vs 11.9 chrF baseline).

Use

This is a custom architecture — not loadable via AutoModel. Clone the repo, drop model.pt in, and use the provided loader:

git clone https://github.com/jakequist/kanjiland && cd kanjiland
uv run python scripts/evaluate.py --config config.yaml --checkpoint model.pt --test-sets kftt-test

Files: model.pt (weights + config), config.yaml, tokenizer-16k.json.

License & provenance

Weights: MIT. Trained on a corpus that is ~86% JParaCrawl (research-use, not redistributed) — publishing weights does not redistribute the corpus. Formal-domain (KFTT) scores trail mixed-domain by ~8 chrF (web-domain training skew).

Downloads last month
10
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support