grc-macronizer-char (seed 17) -- experimental variant

This is one of three random-seed replicates of Ericu950/oga-macronizer-char, provided for experimentation and reproducibility only. Use the main repo unless you specifically need this seed.

Same architecture, training data, and hyperparameters as the main model (see that repo's README for the full description); only the random seed (--seed 17) differs, controlling weight initialization and the train/validation split draw.

Evaluation

Full Norma Syllabarum Graecarum benchmark (17 works, 2,957 gold-annotated positions), scored with macron_model/eval_norma.py:

system accuracy (unmarked defaults to short) raw accuracy unmarked / total
rule-based grc-macronizer 86.84% 53.77% 1,300 / 2,957
this checkpoint (seed 17) 89.72% 89.62% 9 / 2,957
main repo (seed 42) 89.89% 89.79% 9 / 2,957

Across three seeds, defaults-to-short accuracy averages 89.36±0.78% (88.47% / 89.72% / 89.89%), with every seed beating the rule-based teacher -- this checkpoint is the middle of the three.

Usage

import sys
sys.path.insert(0, "path/to/grc-macronizer/macron_model")  # for predict.py
from predict import MacronPredictor

predictor = MacronPredictor("path/to/downloaded/checkpoint")
macronized = predictor.macronize("ανθρωπος ανηρ")

Citation

Thörn Cleland, Albin and Eric Cullhed (forthcoming). Automatic Annotation of Ancient Greek Vowel Length.

License

GNU GPL v3, matching grc-macronizer.

Downloads last month
48
Safetensors
Model size
867k params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support