Ericu950's picture
Upload README.md with huggingface_hub
da47f43 verified
|
Raw
History Blame Contribute Delete
2 kB
---
language: grc
license: gpl-3.0
tags:
- ancient-greek
- vowel-length
- macronization
- token-classification
- char-level
---
# grc-macronizer-char (seed 13, original) -- experimental variant
**This is one of three random-seed replicates of
[Ericu950/oga-macronizer-char](https://huggingface.co/Ericu950/oga-macronizer-char),
provided for experimentation and reproducibility only. Use the main repo
unless you specifically need this seed.**
This was the original single-seed checkpoint (default `--seed 13`, no explicit
override) previously published as the main repo; it has since been superseded
there by a better-performing seed. Same architecture, training data, and
hyperparameters as the main model (see that repo's README for the full
description) -- only the random seed differs.
## Evaluation
Full [Norma Syllabarum Graecarum](https://huggingface.co/datasets/Urdatorn/norma)
benchmark (17 works, 2,957 gold-annotated positions), scored with
`macron_model/eval_norma.py`:
| system | accuracy (unmarked defaults to short) | raw accuracy | unmarked / total |
|---|---|---|---|
| rule-based grc-macronizer | 86.84% | 53.77% | 1,300 / 2,957 |
| this checkpoint (seed 13) | 88.47% | 88.37% | 9 / 2,957 |
| main repo (seed 42) | 89.89% | 89.79% | 9 / 2,957 |
Across three seeds, defaults-to-short accuracy averages 89.36±0.78%
(88.47% / 89.72% / 89.89%), with every seed beating the rule-based teacher --
this checkpoint is the weakest of the three, though still ahead of the
rule-based system.
## Usage
```python
import sys
sys.path.insert(0, "path/to/grc-macronizer/macron_model") # for predict.py
from predict import MacronPredictor
predictor = MacronPredictor("path/to/downloaded/checkpoint")
macronized = predictor.macronize("ανθρωπος ανηρ")
```
## Citation
> Thörn Cleland, Albin and Eric Cullhed (forthcoming). Automatic Annotation of Ancient Greek Vowel Length.
## License
GNU GPL v3, matching [grc-macronizer](https://github.com/Urdatorn/grc-macronizer).