# Reproducing Stoicheia ## 1. Environment ```bash git clone https://github.com/ericu9500/stoicheia cd stoicheia export STOICHEIA_DATA=/path/to/a/writable/data+checkpoints/directory source env.sh pip install -e . ``` For GPU training at scale, the project used an Apptainer container on aarch64 GH200 nodes via a SLURM cluster; the container image itself is not included in this repo (see `scripts/slurm/README.md` for what's expected inside it, and how to point `STOICHEIA_SIF` at your own equivalent, or run without a container at all). See the same file for the cluster-specific sbatch templates (edit the account/partition placeholders before submitting). For CPU-only work (tests, small-scale inference, data preparation), `pip install -r requirements-cpu.txt` is sufficient — no container needed, `attn_impl="sdpa"` runs on CPU. ## 2. Data Three core datasets are already public and used as-is (no re-release needed): ```python from datasets import load_dataset gold_silver = load_dataset("Ericu950/AncientGreek") # pretraining corpus bronze = load_dataset("Ericu950/SyntheticAncientGreek-CorpusCorporum") # synthetic augmentation inscriptions = load_dataset("Ericu950/Inscriptions_2") # PHI inscriptions ``` Two further data dependencies are external, citable resources — clone/download them directly rather than expecting a copy in this repo: - **OGA/AGDT treebank** (tagging/parsing fine-tuning, morphosyntax evaluation): clone Celano's own repository, `git.informatik.uni-leipzig.de/celano/morphosyntactic_parser_for_oga`. - **Norma** (macronization/scansion benchmark): released with the vowel-length paper cited in the paper's macronization section; the harnesses read it through `MACRONIZER_SRC` (see the last section of this file). The 10-fold decontamination split is built by the pipeline in `data/split_pipeline/` (MinHash-LSH near-duplicate clustering, a last-digit rule for papyri/inscriptions, and n-gram decontamination against the eval sets — see the last section of this file and `data/fold_manifests/SPLIT_DESIGN.md`). Its downstream consumer, `data/build_fold_shards.py`, takes a fold's `train.jsonl.zst` and builds the memmap shards the pretraining loader reads. In practice you don't need to rebuild the fold assignments from scratch: use the pretrained checkpoints directly (`MODEL_CARDS_INDEX.md`), and `data/fold_manifests/` records every fold's test-work assignments for verification. ## 3. Pretraining (the 11 backbones) ```bash sbatch scripts/slurm/pretrain_fold.sbatch 0 # one of ten literary folds (0-9) sbatch scripts/slurm/pretrain_doc_clean.sbatch # the documentary-clean model ``` Each is a long-running, checkpointed, resumable job chain (dev-driven schedule, not step-capped — see the paper's Model section for the staged-anneal training regime). Pretrained checkpoints are also available directly on the Hub (see `MODEL_CARDS_INDEX.md`) — you do not need to re-pretrain to use or fine-tune the models. ## 4. Fine-tuning the three downstream tasks All three fine-tune from `$STOICHEIA_DATA/runs/stoicheia_doc_clean/best.pt` (or the equivalent Hub checkpoint, downloaded locally first if you want to fine-tune outside this pipeline's own checkpoint format). ```bash # restoration, v2 recipe, held-out digit 3 (the paper's headline model) sbatch scripts/slurm/insc_finetune_whole_4node.sbatch configs/insc/finetune_whole_v4_t3v4.json # joint tagger + dependency parser sbatch scripts/slurm/syntax_joint_ddp.sbatch configs/syntax/joint_docclean_f3_s0.json # macronization + metrical scansion (joint), and the macron-only arm of the ablation sbatch scripts/slurm/meter_meter.sbatch configs/meter/joint_docclean.json sbatch scripts/slurm/meter_meter.sbatch configs/meter/mac_v2.json ``` For the leak-proof 10-fold restoration rotation used to interrogate individual inscriptions/papyri (every document gets a fine-tuned model that provably never saw it, regardless of which digit its ID ends in — restoration fine-tunes in about an hour, so the full rotation is cheap): ```bash # test_digit/val_digit rotate together: (1,2), (2,3), ..., (9,0), (0,1) sbatch scripts/slurm/insc_finetune_whole_4node.sbatch configs/insc/finetune_whole_v4_t1v2.json ``` Each `configs/insc/finetune_whole_v*_t*v*.json` pins its own `test_digit`/ `val_digit` fields; `insc/train/finetune_whole.py` reads them and exports `INSC_TEST_DIGIT`/`INSC_VAL_DIGIT` itself before any data loads, so the intended split holds however the script is launched. The eval scripts (`insc/eval/restore_strict{,_papyri}.py`) are not config-driven — when evaluating one of these fold models, export the matching `INSC_TEST_DIGIT`/`INSC_VAL_DIGIT` yourself before `--make-samples`/`--ckpt` (defaults to the flagship 3/4 split otherwise). ## 5. Evaluation (reproducing the paper's tables) Every reconstruction number uses one protocol: whole documents, real lacunae left in context as unknowns, spaces counted toward the gap length, word division predicted, beam 20, Levenshtein CER. `--make-samples` writes a frozen sample file; evaluation then reads it, so every system sees identical gaps. **Ten-fold rotation (Table 1).** For each held-out digit `t` (val digit `v = (t+1) % 10`), with the arm's checkpoint from `MODEL_CARDS_INDEX.md`: ```bash export INSC_TEST_DIGIT=3 INSC_VAL_DIGIT=4 # must match the model's fold python -m insc.eval.restore_strict --make-samples --n 100 --samples f3_inscr.json python -m insc.eval.restore_strict_papyri --make-samples --n 100 --samples f3_pap.json python -m insc.eval.restore_strict --ckpt /best.pt --samples f3_inscr.json --out v2_d3_inscr.json python -m insc.eval.restore_strict_papyri --ckpt /best.pt --samples f3_pap.json --out v2_d3_pap.json ``` Repeat over the ten `v4_*` runs (paper revision v2), the ten `v3_*` runs (v1), and the ten `v3_randinit_*` runs (the matched control), then aggregate and run the fold-paired permutation tests with `python -m analysis.sig_all`. **Starting from a released checkpoint.** The Hub ships `model.safetensors` for `AutoModel.from_pretrained`, while the evaluation scripts read the trainer's `.pt`. Convert once: ```bash python scripts/hf_to_eval_checkpoint.py \ --repo Ericu950/Stoicheia-restoration-test3 --out runs/t3v4/best.pt ``` The result is bit-identical to the checkpoint the paper evaluated (268 tensors, max absolute difference 0.0). Export `INSC_TEST_DIGIT`/`INSC_VAL_DIGIT` to match the checkpoint's held-out digit before running any evaluation below. **Head-to-head against Ithaca and Aeneas (Table 2).** All three systems read the frozen 3,000-sample digit-3 file shipped in `insc/eval/frozen/`: ```bash python -m insc.eval.restore_strict --ckpt /best.pt \ --samples insc/eval/frozen/strict_test_fold3_samples.json --out ours_strict.json python -m insc.eval.ithaca_baseline --samples insc/eval/frozen/strict_test_fold3_samples.json --out ithaca.json python -m insc.eval.ptp_baseline --ckpt .pkl \ --samples insc/eval/frozen/strict_test_fold3_samples.json --out aeneas.json ``` **Recently edited documents (Tables 3-4).** Scored on the comparison release's own documents and normalization, with every system rescored under the same Levenshtein CER: ```bash python -m insc.eval.restore_dsh --ckpt /best.pt --data .jsonl --out ours_recent.json python -m insc.eval.ptp_baseline --ckpt .pkl --dsh .jsonl --out aeneas_recent.json python -m analysis.merge_dsh # aggregate shards, macro over gap lengths ``` **Tagging and parsing (Table 5).** The full 5-fold x 2-seed matrix per encoder: ```bash for f in 0 1 2 3 4; do for s in 0 1; do sbatch scripts/slurm/syntax_joint_ddp.sbatch configs/syntax/joint_docclean_f${f}_s${s}.json done; done python -m parser.joint_evaluate --run $STOICHEIA_DATA/parser_data/runs/joint_docclean_f0_s0 --split test ``` **Macronization and scansion (Tables 6-7).** The macronization ablation uses the macron-only runs (`mac_v2*`, six seeds per arm), the scansion ablation the joint runs (`joint_docclean*` / `joint_randinit*`); the external macronizer comparison scores the joint model on Norma's 1,916 test positions: ```bash python -m meter.predict --model $STOICHEIA_DATA/runs/meter_joint_docclean/best.pt --norma python -m meter.predict --model $STOICHEIA_DATA/runs/meter_mac_v2/best.pt --norma --norma-source git ``` `python -m analysis.paper_tables` collects finished evaluation logs into the table layouts used in the paper. ## 6. Tests ```bash pytest tests/ -q ``` CPU-only, seconds-scale. Covers normalization/packing/noising (pretraining), edit-script lemma encoding + dataset construction (tagger), macron/scansion mark parsing (meter), plus lightweight forward-pass smoke tests for the restoration and joint tagger/parser pipelines (tiny randomly-initialized configs — these check the tensor plumbing survives refactors, not model quality). ## Reproducing the 10-fold split itself The split is not an input -- it is constructed by the 13-stage pipeline in `data/split_pipeline/` (sentence segmentation, pristine-edition clustering via exact/MinHash-LSH/shared-sentence evidence, zone assignment, per-fold exclusion masks, documentary decontamination incl. blocklist, bronze back-translation filters, verification, manifests). `data/fold_manifests/SPLIT_DESIGN.md` documents the design; `fold_k_test_works.tsv.gz` lists every work in fold k's test bucket (group id, kind, source, record/char counts, author+title), so the "provably never seen" guarantee is checkable work-by-work without rebuilding anything. ## Revision naming The paper's reconstruction revisions v1 and v2 correspond, for historical reasons, to configuration files named `finetune_whole_v3_*` and `finetune_whole_v4_*` respectively (earlier internal iterations v1/v2 were superseded before evaluation and are not part of the release). ## External pieces the harnesses expect * Meter training data: `MACRONIZER_SRC` must point at a checkout of the `Ericu950/Stoicheia-meter-silver` dataset (silver verse lines + scanner corpus); the Norma benchmark itself comes from `Ericu950/norma`. * The Ithaca baseline harness (`insc/eval/ithaca_baseline.py`) expects DeepMind's Ithaca repository cloned at `ithaca_upstream/` and its released checkpoint; the Aeneas harness (`insc/eval/ptp_baseline.py`) expects DeepMind's `predictingthepast` repository and its Greek checkpoint. Both are public third-party releases and are not vendored here.