--- language: - ang license: mit library_name: pytorch tags: - lemmatization - old-english - anglo-saxon - historical-nlp - low-resource - morphology - transformer - character-level pipeline_tag: token-classification model-index: - name: old-english-lemmatizer results: - task: type: token-classification name: Lemmatisation dataset: name: YCOE (corpus-only held-out test split, 80/10/10, seed 42) type: ycoe metrics: - type: accuracy value: 67.58 name: Test lemma accuracy - type: accuracy value: 82.75 name: Test POS accuracy --- # Model Card for old-english-lemmatizer A character-level Transformer encoder–decoder that lemmatises **Old English** (c. 450–1150 CE) and jointly predicts a Universal Dependencies (UPOS) tag. It maps the character sequence of an inflected form to the character sequence of its lemma, processing words in isolation (context-free). This is the best (improved) model behind the two companion publications, reaching **67.58 % lemma accuracy** and **82.75 % POS accuracy** on the held-out corpus-only test set. ## Model Details ### Model Description - **Developed by:** Javier Martín Arista (Nerthus Project, Universidad de La Rioja) - **Funded by:** Grant PID2023-149762NB-100 (Agencia Estatal de Investigación); grant PRXX24/00108 (Ministerio de Ciencia, Innovación y Universidades) - **Model type:** Character-level Transformer encoder–decoder with a joint UPOS classification head (sequence-to-sequence character transduction) - **Language(s):** Old English (`ang`) - **License:** MIT (source code and released weights only; not the training corpora) - **Reference version:** Where the two companion papers differ, the conference paper (LaTeLL 2026) is authoritative. ### Model Sources - **Repository:** https://huggingface.co/Nerthus-Project/old-english-lemmatizer - **Paper (conference, reference):** Martín Arista, J. (2026). *Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation.* LaTeLL 2026. - **Paper (journal):** Martín Arista, J. (2026). *Neural Architectures for Old English Lemmatisation: A Comparative Study of Corpus-Based and Dictionary-Augmented Approaches.* Procesamiento del Lenguaje Natural (SEPLN). ## Uses ### Direct Use Lemmatisation and UPOS tagging of individual Old English word forms for corpus linguistics and digital-humanities workflows: concordance generation, frequency analysis, and dictionary linking. Input is an inflected form; output is `(lemma, UPOS)`. ### Downstream Use As a preprocessing component feeding search/concordancing tools, editorial workflows, or lexical-resource linking over digitised Old English texts. ### Out-of-Scope Use - Sentence-level or context-dependent disambiguation of homographs (the model is context-free; e.g. it cannot resolve *þā* demonstrative vs. adverb from context). - Languages or historical stages other than Old English. - Named-entity-aware lemmatisation (proper nouns are the weakest category). ## Bias, Risks, and Limitations - Processes words in isolation; no syntactic context. - Errors concentrate in strong-verb ablaut (~25 %), spelling variation (~20 %), and prefix (`ge-`) handling (~15 %); proper nouns and hapax legomena are hardest. - The gold data contain ~11.3 % inconsistently lemmatised word types, which caps achievable accuracy (human agreement ≈ 90–95 %). - Evaluated only on a held-out split of the YCOE; behaviour on independently collected texts or other editions may differ. ### Recommendations Treat outputs as high-quality suggestions to be validated for ambiguous forms, proper nouns, and rare words. For homographs requiring context, add post-processing or human review. ## How to Get Started with the Model The weights (`model.pt`) are a raw PyTorch `state_dict`; load them with the bundled code (this is **not** a `transformers` `AutoModel`). ```python # pip install -r requirements.txt (torch>=2.0) from lemmatizer import OldEnglishLemmatizer lem = OldEnglishLemmatizer.from_pretrained(".") # folder with config.json, char_vocab.json, model.pt lem.lemmatize("cyninges") # ('cyning', 'NOUN') lem.lemmatize_batch(["cwæð", "folce", "cyning"]) # [('cweþan', 'VERB'), ('folc', 'NOUN'), ('cyning', 'NOUN')] lem.lemmatize_text("se cyning cwæð to his folce") # [{'word': 'se', 'lemma': ..., 'pos': ...}, ...] ``` ## Training Details ### Training Data Trained on the **corpus-only** condition: word–lemma–POS triples extracted from the York–Toronto–Helsinki Parsed Corpus of Old English (YCOE, prose + poetry). The dictionaries and corpora **cannot be redistributed** and are not included here; obtain them from their official providers: | Resource | Role | Official source | |---|---|---| | YCOE Prose (Taylor et al. 2003) | Training corpus (prose) | https://www-users.york.ac.uk/~lang22/YCOE/YcoeHome.htm | | York–Helsinki Poetry (Pintzuk & Plug 2001) | Training corpus (poetry) | https://www-users.york.ac.uk/~lang18/pcorpus.html | | DOE: A to I (Healey 2018) | Dictionary augmentation (ablation) | https://www.doe.utoronto.ca/ | | DOEC (Healey 2012) | Word2Vec pre-training (ablation) | https://www.doe.utoronto.ca/pages/pub/web-corpus.html | | Bosworth–Toller (Crist & Tichý 2014) | Dictionary augmentation (ablation) | https://bosworthtoller.com | | ParCorOEv3 (Martín Arista et al. 2023) | Related UD resource | https://huggingface.co/Nerthus-Project | | VARIOE (Cichosz et al. 2021) | Related lemma inventory | http://varioe.pelcra.pl/morph | Dataset statistics (corpus-only, after harmonisation): 105,063 annotated instances; 84,050 / 10,506 / 10,507 train/dev/test (80/10/10, seed 42); 90,752 unique word types; 37,627 unique lemmas; 28.3 % identity transformations; 4.4 % unknown POS; 11.7 % `ge-` prefix; 11.3 % inconsistent types. ### Training Procedure #### Preprocessing Extract triples from parsed trees (prose columns LEMMA–WORD–POS, poetry WORD–LEMMA–POS); map all POS labels to UPOS (unresolved → `X`); Unicode-normalise and case-fold while preserving macrons; normalise 221 parenthetical `(ge)` lemmas; remove exact-duplicate triples; 80/10/10 random split with seed 42. #### Training Hyperparameters - **Training regime:** fp32 - **Epochs:** 80 (best at epoch 74) - **Batch size:** 24 (effective 72, gradient accumulation 3) - **Optimizer:** AdamW, weight decay 0.01 - **Learning rate:** 1.5e-4, cosine annealing with warm restarts (range 1e-5–1.5e-4, 4 cycles, peaks at epochs 22/41/58/74), 5 % warmup - **Label smoothing:** 0.15 - **POS loss weight:** 0.2 - **Data augmentation:** character-level (substitution `þ↔ð`, `æ↔a/e`, `i↔y`; deletion; adjacent swap; duplication), p=0.2 per word, re-applied each epoch - **Seed:** 42 ## Evaluation ### Testing Data, Factors & Metrics - **Testing data:** YCOE corpus-only held-out test split (10,507 instances). - **Metrics:** exact-match lemma accuracy (after normalisation) and exact-match UPOS accuracy. ### Results | Model | Dev lemma | Test lemma | Test POS | |---|---|---|---| | **This model** (Transformer + augmentation + cyclical LR) | **68.68 %** | **67.58 %** | **82.75 %** | | Transformer baseline | 66.65 % | — | — | | Seq2Seq (BiLSTM + attention) baseline | 53.15 % | — | — | Ablations (test accuracy): Word2Vec context fusion 63.53 % (−4.05); Word2Vec + character concat 49.37 % (−12.67); BPE 500-merge 54.66 % (−12.92). Approximate accuracy by POS: ADP ~85 %, CCONJ ~83 %, DET ~80 %, ADV ~75 %, NOUN ~68 %, ADJ ~65 %, VERB ~55 %, PROPN ~45 %. #### Summary A corpus-only, character-level Transformer with morphology-aware augmentation and cyclical learning rates gives the best reported accuracy for Old English lemmatisation on the YCOE; word embeddings and subword tokenisation both hurt. ## Technical Specifications ### Model Architecture and Objective Character-level Transformer encoder–decoder (`d_model` 512, 8 heads, 4 encoder + 4 decoder layers, feed-forward 1024, dropout 0.15, sinusoidal positional encoding) with an auxiliary UPOS classification head. Character vocabulary of 140 symbols (incl. `þ ð æ ƿ`, macron vowels, Latin script, digits, punctuation, and 4 special tokens); 18 UPOS tags; ≈ 26.5 M parameters. Objective: character-level cross-entropy for the lemma plus cross-entropy for the UPOS tag. Released weights are used in single-word (context-free) mode. > **Reconciliation note.** The papers describe the model with a "98-symbol" > vocabulary and ≈ 21.5 M parameters (the plain Transformer baseline). The > released best checkpoint ships a 140-entry operational vocabulary and has > ≈ 26.5 M parameters. Evaluation figures are unaffected. See > `NOTES_ON_PUBLICATIONS.md`. ### Compute Infrastructure - **Software:** Python ≥ 3.9, PyTorch ≥ 2.0. ## Citation Please cite **both** works. **BibTeX:** ```bibtex @inproceedings{MartinArista2026LaTeLL, title = {Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation}, author = {Mart{\'i}n Arista, Javier}, booktitle = {Proceedings of LaTeLL 2026 (Workshop on Language Technologies for Historical and Ancient Languages)}, year = {2026} } @article{MartinArista2026PLN, title = {Neural Architectures for Old English Lemmatisation: A Comparative Study of Corpus-Based and Dictionary-Augmented Approaches}, author = {Mart{\'i}n Arista, Javier}, journal = {Procesamiento del Lenguaje Natural}, year = {2026} } ``` ## Model Card Authors Javier Martín Arista (Nerthus Project, Universidad de La Rioja). ## Model Card Contact javier.martin@unirioja.es