| --- |
| license: mit |
| language: [ja, en] |
| tags: [japanese, reading-comprehension, annotation, furigana, grammar, from-scratch] |
| --- |
| # Kanjiland — Annotation model (Ja→full format) |
|
|
| A from-scratch 52.3M Transformer that maps raw Japanese to the **full Kanjiland |
| reading-comprehension format** — segmentation, furigana, glosses, word groupings, |
| translation, and grammar labels (⟨T⟩/⟨W⟩/⟨S⟩/⟨G⟩), with **no MeCab at inference** |
| (it learned segmentation itself). Part of |
| [**Kanjiland**](https://github.com/jakequist/kanjiland). |
|
|
| - **Training:** 6,445 teacher-supervised silver annotations (the M7 dataset). |
| - **Format-validity eval:** parse-rate **77%**, fully-valid (linter) **38%**. |
| - **On-device:** runs on CPU at 521 tok/s; int8 is 2.4× smaller. |
|
|
| ## ⚠ Numeric fragility |
| Generation **must** run under bf16 autocast (cpu + cuda). In fp32 the long |
| autoregressive decode diverges to 0% parseable output. The loader handles this. |
|
|
| ## Use |
| ```bash |
| git clone https://github.com/jakequist/kanjiland && cd kanjiland |
| uv run python scripts/annotate.py --config config.yaml --checkpoint model.pt \ |
| --text "彼は古い寺を訪れた。" --device cpu |
| ``` |
|
|
| ## License & limitations |
| Weights: **MIT**. This is an **early de-risk baseline** trained on only 6.8k |
| examples — valid *structure*, rough *content* (glosses/translations loop). Improve |
| via more silver data + constrained decoding, not architecture. Grammar inventory: |
| [docs/GRAMMAR_RULES.md](https://github.com/jakequist/kanjiland/blob/main/docs/GRAMMAR_RULES.md). |
|
|