File size: 4,348 Bytes
8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d fb071da 40c7be8 8b4c84d 40c7be8 fb071da 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 8b4c84d 40c7be8 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 | ---
language: grc
license: gpl-3.0
tags:
- ancient-greek
- vowel-length
- macronization
- token-classification
- char-level
---
# grc-macronizer-char: a small char-level model for Ancient Greek vowel-length annotation
This model predicts the phonemic length (long/short/undetermined) of the three
Ancient Greek "dichrona" -- alpha (α), iota (ι), and upsilon (υ) -- whose length
is not disambiguated by the standard Greek script. It was trained on Ancient
Greek text automatically macronized by the rule-based
[grc-macronizer](https://github.com/Urdatorn/grc-macronizer), applied to the
[Opera Graeca Adnotata](https://doi.org/10.5281/zenodo.14206061) corpus
(Celano 2024), and is intended as a complementary, corpus-general alternative
to that rule-based system.
## Architecture
A small transformer encoder (~0.9M parameters) operating over three character-level
"planes" per input position:
- **plane 1 (letter)**: which of the 24 Greek letters (final/medial sigma
folded together), or space, or "other" (punctuation/digits/foreign chars)
- **plane 2 (diacritic)**: the combination of accent/breathing/diaeresis marks
on that letter, if any (a small vocabulary fit from the training data)
- **target (macron)**: none / short / long -- predicted only at positions
that are genuine ambiguous dichrona (diphthong members and circumflexed
vowels are excluded, matching the rule-based system's own definition)
During training, the diacritic plane is randomly masked (accents stripped)
for a fraction of characters, so the model also learns to macronize
unaccented input.
## Training data
Training labels come directly from the rule-based macronizer's own output:
wherever it marks a dichronon (it does not guess -- see the
[macronizer paper](https://github.com/Urdatorn/grc-macronizer)), that becomes
a training label; dichrona it left unmarked are excluded from the loss
entirely (never treated as a negative "short" or "long" example), since the
rule-based system is known to be incomplete rather than wrong where it commits.
## Evaluation
Evaluated on the full [Norma Syllabarum Graecarum](https://huggingface.co/datasets/Urdatorn/norma)
benchmark, spanning all 17 manually-annotated works (prose and verse; 2,957
gold-annotated positions). Scored with the identical harness
(`macron_model/eval_norma.py`) against the bundled rule-based macronizer,
counting a left-unmarked dichronon as an implicit "short" guess for both
systems (`default_short_accuracy` below) so abstention isn't unfairly
penalized either way:
| system | accuracy (unmarked defaults to short) | raw accuracy | unmarked / total |
|---|---|---|---|
| rule-based grc-macronizer | 86.84% | 53.77% | 1,300 / 2,957 |
| **this model** | **89.89%** | **89.79%** | **9 / 2,957** |
The model outperforms the rule-based system on the fair metric while resolving
essentially every position (99.7% coverage vs. 56%), since it learned to
generalize past the cases the rule-based system couldn't commit to. The margin
is statistically significant (McNemar's test, continuity-corrected,
χ²=10.04, p=0.0015).
Two additional training runs with different random seeds confirm this is not
a lucky draw: defaults-to-short accuracy across three seeds averages
89.36±0.78% (88.47% / 89.72% / **89.89%**, this checkpoint), with every
seed beating the rule-based teacher.
## Usage
```python
import sys
sys.path.insert(0, "path/to/grc-macronizer/macron_model") # for predict.py
from predict import MacronPredictor
predictor = MacronPredictor("path/to/downloaded/checkpoint")
macronized = predictor.macronize("ανθρωπος ανηρ")
# -> "α^νθρωπος α_νηρ"
```
`MacronPredictor` expects the checkpoint directory (containing
`config.json`/`model.safetensors`) plus a sibling or parent `diacritic_vocab.json`
(see `predict.py` for the exact lookup logic). The `predict.py`/`tokenizer.py`
source lives in [grc-macronizer/macron_model](https://github.com/Urdatorn/grc-macronizer/tree/oga-macronization-improvements/macron_model).
## Citation
If you use this model, please cite:
> Thörn Cleland, Albin and Eric Cullhed (forthcoming). Automatic Annotation of Ancient Greek Vowel Length.
## License
GNU GPL v3, matching [grc-macronizer](https://github.com/Urdatorn/grc-macronizer)
(the training-data generator this model's code ships alongside).
|