jamartina's picture
Update README.md
9702d6c verified
|
Raw
History Blame Contribute Delete
9.66 kB
---
language:
- ang
license: mit
library_name: pytorch
tags:
- lemmatization
- old-english
- anglo-saxon
- historical-nlp
- low-resource
- morphology
- transformer
- character-level
pipeline_tag: token-classification
model-index:
- name: old-english-lemmatizer
results:
- task:
type: token-classification
name: Lemmatisation
dataset:
name: YCOE (corpus-only held-out test split, 80/10/10, seed 42)
type: ycoe
metrics:
- type: accuracy
value: 67.58
name: Test lemma accuracy
- type: accuracy
value: 82.75
name: Test POS accuracy
---
# Model Card for old-english-lemmatizer
A character-level Transformer encoder–decoder that lemmatises **Old English**
(c. 450–1150 CE) and jointly predicts a Universal Dependencies (UPOS) tag. It
maps the character sequence of an inflected form to the character sequence of its
lemma, processing words in isolation (context-free). This is the best (improved)
model behind the two companion publications, reaching **67.58 % lemma accuracy**
and **82.75 % POS accuracy** on the held-out corpus-only test set.
## Model Details
### Model Description
- **Developed by:** Javier Martín Arista (Nerthus Project, Universidad de La Rioja)
- **Funded by:** Grant PID2023-149762NB-100 (Agencia Estatal de Investigación); grant PRXX24/00108 (Ministerio de Ciencia, Innovación y Universidades)
- **Model type:** Character-level Transformer encoder–decoder with a joint UPOS classification head (sequence-to-sequence character transduction)
- **Language(s):** Old English (`ang`)
- **License:** MIT (source code and released weights only; not the training corpora)
- **Reference version:** Where the two companion papers differ, the conference paper (LaTeLL 2026) is authoritative.
### Model Sources
- **Repository:** https://huggingface.co/Nerthus-Project/old-english-lemmatizer
- **Paper (conference, reference):** Martín Arista, J. (2026). *Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation.* LaTeLL 2026.
- **Paper (journal):** Martín Arista, J. (2026). *Neural Architectures for Old English Lemmatisation: A Comparative Study of Corpus-Based and Dictionary-Augmented Approaches.* Procesamiento del Lenguaje Natural (SEPLN).
## Uses
### Direct Use
Lemmatisation and UPOS tagging of individual Old English word forms for corpus
linguistics and digital-humanities workflows: concordance generation, frequency
analysis, and dictionary linking. Input is an inflected form; output is
`(lemma, UPOS)`.
### Downstream Use
As a preprocessing component feeding search/concordancing tools, editorial
workflows, or lexical-resource linking over digitised Old English texts.
### Out-of-Scope Use
- Sentence-level or context-dependent disambiguation of homographs (the model is
context-free; e.g. it cannot resolve *þā* demonstrative vs. adverb from context).
- Languages or historical stages other than Old English.
- Named-entity-aware lemmatisation (proper nouns are the weakest category).
## Bias, Risks, and Limitations
- Processes words in isolation; no syntactic context.
- Errors concentrate in strong-verb ablaut (~25 %), spelling variation (~20 %),
and prefix (`ge-`) handling (~15 %); proper nouns and hapax legomena are hardest.
- The gold data contain ~11.3 % inconsistently lemmatised word types, which caps
achievable accuracy (human agreement ≈ 90–95 %).
- Evaluated only on a held-out split of the YCOE; behaviour on independently
collected texts or other editions may differ.
### Recommendations
Treat outputs as high-quality suggestions to be validated for ambiguous forms,
proper nouns, and rare words. For homographs requiring context, add
post-processing or human review.
## How to Get Started with the Model
The weights (`model.pt`) are a raw PyTorch `state_dict`; load them with the
bundled code (this is **not** a `transformers` `AutoModel`).
```python
# pip install -r requirements.txt (torch>=2.0)
from lemmatizer import OldEnglishLemmatizer
lem = OldEnglishLemmatizer.from_pretrained(".") # folder with config.json, char_vocab.json, model.pt
lem.lemmatize("cyninges")
# ('cyning', 'NOUN')
lem.lemmatize_batch(["cwæð", "folce", "cyning"])
# [('cweþan', 'VERB'), ('folc', 'NOUN'), ('cyning', 'NOUN')]
lem.lemmatize_text("se cyning cwæð to his folce")
# [{'word': 'se', 'lemma': ..., 'pos': ...}, ...]
```
## Training Details
### Training Data
Trained on the **corpus-only** condition: word–lemma–POS triples extracted from
the York–Toronto–Helsinki Parsed Corpus of Old English (YCOE, prose + poetry).
The dictionaries and corpora **cannot be redistributed** and are not included
here; obtain them from their official providers:
| Resource | Role | Official source |
|---|---|---|
| YCOE Prose (Taylor et al. 2003) | Training corpus (prose) | https://www-users.york.ac.uk/~lang22/YCOE/YcoeHome.htm |
| York–Helsinki Poetry (Pintzuk & Plug 2001) | Training corpus (poetry) | https://www-users.york.ac.uk/~lang18/pcorpus.html |
| DOE: A to I (Healey 2018) | Dictionary augmentation (ablation) | https://www.doe.utoronto.ca/ |
| DOEC (Healey 2012) | Word2Vec pre-training (ablation) | https://www.doe.utoronto.ca/pages/pub/web-corpus.html |
| Bosworth–Toller (Crist & Tichý 2014) | Dictionary augmentation (ablation) | https://bosworthtoller.com |
| ParCorOEv3 (Martín Arista et al. 2023) | Related UD resource | https://huggingface.co/Nerthus-Project |
| VARIOE (Cichosz et al. 2021) | Related lemma inventory | http://varioe.pelcra.pl/morph |
Dataset statistics (corpus-only, after harmonisation): 105,063 annotated
instances; 84,050 / 10,506 / 10,507 train/dev/test (80/10/10, seed 42); 90,752
unique word types; 37,627 unique lemmas; 28.3 % identity transformations; 4.4 %
unknown POS; 11.7 % `ge-` prefix; 11.3 % inconsistent types.
### Training Procedure
#### Preprocessing
Extract triples from parsed trees (prose columns LEMMA–WORD–POS, poetry
WORD–LEMMA–POS); map all POS labels to UPOS (unresolved → `X`); Unicode-normalise
and case-fold while preserving macrons; normalise 221 parenthetical `(ge)`
lemmas; remove exact-duplicate triples; 80/10/10 random split with seed 42.
#### Training Hyperparameters
- **Training regime:** fp32
- **Epochs:** 80 (best at epoch 74)
- **Batch size:** 24 (effective 72, gradient accumulation 3)
- **Optimizer:** AdamW, weight decay 0.01
- **Learning rate:** 1.5e-4, cosine annealing with warm restarts (range 1e-5–1.5e-4, 4 cycles, peaks at epochs 22/41/58/74), 5 % warmup
- **Label smoothing:** 0.15
- **POS loss weight:** 0.2
- **Data augmentation:** character-level (substitution `þ↔ð`, `æ↔a/e`, `i↔y`; deletion; adjacent swap; duplication), p=0.2 per word, re-applied each epoch
- **Seed:** 42
## Evaluation
### Testing Data, Factors & Metrics
- **Testing data:** YCOE corpus-only held-out test split (10,507 instances).
- **Metrics:** exact-match lemma accuracy (after normalisation) and exact-match UPOS accuracy.
### Results
| Model | Dev lemma | Test lemma | Test POS |
|---|---|---|---|
| **This model** (Transformer + augmentation + cyclical LR) | **68.68 %** | **67.58 %** | **82.75 %** |
| Transformer baseline | 66.65 % | — | — |
| Seq2Seq (BiLSTM + attention) baseline | 53.15 % | — | — |
Ablations (test accuracy): Word2Vec context fusion 63.53 % (−4.05); Word2Vec +
character concat 49.37 % (−12.67); BPE 500-merge 54.66 % (−12.92).
Approximate accuracy by POS: ADP ~85 %, CCONJ ~83 %, DET ~80 %, ADV ~75 %,
NOUN ~68 %, ADJ ~65 %, VERB ~55 %, PROPN ~45 %.
#### Summary
A corpus-only, character-level Transformer with morphology-aware augmentation and
cyclical learning rates gives the best reported accuracy for Old English
lemmatisation on the YCOE; word embeddings and subword tokenisation both hurt.
## Technical Specifications
### Model Architecture and Objective
Character-level Transformer encoder–decoder (`d_model` 512, 8 heads, 4 encoder +
4 decoder layers, feed-forward 1024, dropout 0.15, sinusoidal positional
encoding) with an auxiliary UPOS classification head. Character vocabulary of 140
symbols (incl. `þ ð æ ƿ`, macron vowels, Latin script, digits, punctuation, and 4
special tokens); 18 UPOS tags; ≈ 26.5 M parameters. Objective: character-level
cross-entropy for the lemma plus cross-entropy for the UPOS tag. Released weights
are used in single-word (context-free) mode.
> **Reconciliation note.** The papers describe the model with a "98-symbol"
> vocabulary and ≈ 21.5 M parameters (the plain Transformer baseline). The
> released best checkpoint ships a 140-entry operational vocabulary and has
> ≈ 26.5 M parameters. Evaluation figures are unaffected. See
> `NOTES_ON_PUBLICATIONS.md`.
### Compute Infrastructure
- **Software:** Python ≥ 3.9, PyTorch ≥ 2.0.
## Citation
Please cite **both** works.
**BibTeX:**
```bibtex
@inproceedings{MartinArista2026LaTeLL,
title = {Data Curation, Annotation Quality, and Error Patterns in Old English Automatic Lemmatisation},
author = {Mart{\'i}n Arista, Javier},
booktitle = {Proceedings of LaTeLL 2026 (Workshop on Language Technologies for Historical and Ancient Languages)},
year = {2026}
}
@article{MartinArista2026PLN,
title = {Neural Architectures for Old English Lemmatisation: A Comparative Study of Corpus-Based and Dictionary-Augmented Approaches},
author = {Mart{\'i}n Arista, Javier},
journal = {Procesamiento del Lenguaje Natural},
year = {2026}
}
```
## Model Card Authors
Javier Martín Arista (Nerthus Project, Universidad de La Rioja).
## Model Card Contact
javier.martin@unirioja.es