File size: 10,588 Bytes
5952424 8ce648b 5952424 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 | # Reproducing Stoicheia
## 1. Environment
```bash
git clone https://github.com/ericu9500/stoicheia
cd stoicheia
export STOICHEIA_DATA=/path/to/a/writable/data+checkpoints/directory
source env.sh
pip install -e .
```
For GPU training at scale, the project used an Apptainer container on aarch64 GH200
nodes via a SLURM cluster; the container image itself is not included in this repo (see
`scripts/slurm/README.md` for what's expected inside it, and how to point `STOICHEIA_SIF`
at your own equivalent, or run without a container at all). See the same file for the
cluster-specific sbatch templates (edit the account/partition placeholders before
submitting). For CPU-only work (tests, small-scale inference, data preparation),
`pip install -r requirements-cpu.txt` is sufficient — no container needed,
`attn_impl="sdpa"` runs on CPU.
## 2. Data
Three core datasets are already public and used as-is (no re-release needed):
```python
from datasets import load_dataset
gold_silver = load_dataset("Ericu950/AncientGreek") # pretraining corpus
bronze = load_dataset("Ericu950/SyntheticAncientGreek-CorpusCorporum") # synthetic augmentation
inscriptions = load_dataset("Ericu950/Inscriptions_2") # PHI inscriptions
```
Two further data dependencies are external, citable resources — clone/download them
directly rather than expecting a copy in this repo:
- **OGA/AGDT treebank** (tagging/parsing fine-tuning, morphosyntax evaluation): clone
Celano's own repository, `git.informatik.uni-leipzig.de/celano/morphosyntactic_parser_for_oga`.
- **Norma** (macronization/scansion benchmark): released with the vowel-length paper
cited in the paper's macronization section; the harnesses read it through
`MACRONIZER_SRC` (see the last section of this file).
The 10-fold decontamination split is built by the pipeline in `data/split_pipeline/`
(MinHash-LSH near-duplicate clustering, a last-digit rule for papyri/inscriptions,
and n-gram decontamination against the eval sets — see the last section of this file
and `data/fold_manifests/SPLIT_DESIGN.md`). Its downstream consumer,
`data/build_fold_shards.py`, takes a fold's `train.jsonl.zst` and builds the memmap
shards the pretraining loader reads. In practice you don't need to rebuild the fold
assignments from scratch: use the pretrained checkpoints directly
(`MODEL_CARDS_INDEX.md`), and `data/fold_manifests/` records every fold's test-work
assignments for verification.
## 3. Pretraining (the 11 backbones)
```bash
sbatch scripts/slurm/pretrain_fold.sbatch 0 # one of ten literary folds (0-9)
sbatch scripts/slurm/pretrain_doc_clean.sbatch # the documentary-clean model
```
Each is a long-running, checkpointed, resumable job chain (dev-driven schedule, not
step-capped — see the paper's Model section for the staged-anneal training regime).
Pretrained checkpoints are also available directly on the Hub (see
`MODEL_CARDS_INDEX.md`) — you do not need to re-pretrain to use or fine-tune the models.
## 4. Fine-tuning the three downstream tasks
All three fine-tune from `$STOICHEIA_DATA/runs/stoicheia_doc_clean/best.pt` (or the equivalent
Hub checkpoint, downloaded locally first if you want to fine-tune outside this
pipeline's own checkpoint format).
```bash
# restoration, v2 recipe, held-out digit 3 (the paper's headline model)
sbatch scripts/slurm/insc_finetune_whole_4node.sbatch configs/insc/finetune_whole_v4_t3v4.json
# joint tagger + dependency parser
sbatch scripts/slurm/syntax_joint_ddp.sbatch configs/syntax/joint_docclean_f3_s0.json
# macronization + metrical scansion (joint), and the macron-only arm of the ablation
sbatch scripts/slurm/meter_meter.sbatch configs/meter/joint_docclean.json
sbatch scripts/slurm/meter_meter.sbatch configs/meter/mac_v2.json
```
For the leak-proof 10-fold restoration rotation used to interrogate individual
inscriptions/papyri (every document gets a fine-tuned model that provably never saw it,
regardless of which digit its ID ends in — restoration fine-tunes in about an hour, so
the full rotation is cheap):
```bash
# test_digit/val_digit rotate together: (1,2), (2,3), ..., (9,0), (0,1)
sbatch scripts/slurm/insc_finetune_whole_4node.sbatch configs/insc/finetune_whole_v4_t1v2.json
```
Each `configs/insc/finetune_whole_v*_t*v*.json` pins its own `test_digit`/
`val_digit` fields; `insc/train/finetune_whole.py` reads them and exports
`INSC_TEST_DIGIT`/`INSC_VAL_DIGIT` itself before any data loads, so the intended split
holds however the script is launched. The eval scripts
(`insc/eval/restore_strict{,_papyri}.py`) are not config-driven — when evaluating one of
these fold models, export the matching `INSC_TEST_DIGIT`/`INSC_VAL_DIGIT` yourself
before `--make-samples`/`--ckpt` (defaults to the flagship 3/4 split otherwise).
## 5. Evaluation (reproducing the paper's tables)
Every reconstruction number uses one protocol: whole documents, real lacunae left in
context as unknowns, spaces counted toward the gap length, word division predicted, beam
20, Levenshtein CER. `--make-samples` writes a frozen sample file; evaluation then reads
it, so every system sees identical gaps.
**Ten-fold rotation (Table 1).** For each held-out digit `t` (val digit `v = (t+1) % 10`),
with the arm's checkpoint from `MODEL_CARDS_INDEX.md`:
```bash
export INSC_TEST_DIGIT=3 INSC_VAL_DIGIT=4 # must match the model's fold
python -m insc.eval.restore_strict --make-samples --n 100 --samples f3_inscr.json
python -m insc.eval.restore_strict_papyri --make-samples --n 100 --samples f3_pap.json
python -m insc.eval.restore_strict --ckpt <run>/best.pt --samples f3_inscr.json --out v2_d3_inscr.json
python -m insc.eval.restore_strict_papyri --ckpt <run>/best.pt --samples f3_pap.json --out v2_d3_pap.json
```
Repeat over the ten `v4_*` runs (paper revision v2), the ten `v3_*` runs (v1), and the ten
`v3_randinit_*` runs (the matched control), then aggregate and run the fold-paired
permutation tests with `python -m analysis.sig_all`.
**Starting from a released checkpoint.** The Hub ships `model.safetensors` for
`AutoModel.from_pretrained`, while the evaluation scripts read the trainer's `.pt`. Convert once:
```bash
python scripts/hf_to_eval_checkpoint.py \
--repo Ericu950/Stoicheia-restoration-test3 --out runs/t3v4/best.pt
```
The result is bit-identical to the checkpoint the paper evaluated (268 tensors, max absolute
difference 0.0). Export `INSC_TEST_DIGIT`/`INSC_VAL_DIGIT` to match the checkpoint's held-out
digit before running any evaluation below.
**Head-to-head against Ithaca and Aeneas (Table 2).** All three systems read the frozen
3,000-sample digit-3 file shipped in `insc/eval/frozen/`:
```bash
python -m insc.eval.restore_strict --ckpt <digit-3 run>/best.pt \
--samples insc/eval/frozen/strict_test_fold3_samples.json --out ours_strict.json
python -m insc.eval.ithaca_baseline --samples insc/eval/frozen/strict_test_fold3_samples.json --out ithaca.json
python -m insc.eval.ptp_baseline --ckpt <aeneas>.pkl \
--samples insc/eval/frozen/strict_test_fold3_samples.json --out aeneas.json
```
**Recently edited documents (Tables 3-4).** Scored on the comparison release's own
documents and normalization, with every system rescored under the same Levenshtein CER:
```bash
python -m insc.eval.restore_dsh --ckpt <run>/best.pt --data <recent-inscriptions>.jsonl --out ours_recent.json
python -m insc.eval.ptp_baseline --ckpt <aeneas>.pkl --dsh <recent-inscriptions>.jsonl --out aeneas_recent.json
python -m analysis.merge_dsh # aggregate shards, macro over gap lengths
```
**Tagging and parsing (Table 5).** The full 5-fold x 2-seed matrix per encoder:
```bash
for f in 0 1 2 3 4; do for s in 0 1; do
sbatch scripts/slurm/syntax_joint_ddp.sbatch configs/syntax/joint_docclean_f${f}_s${s}.json
done; done
python -m parser.joint_evaluate --run $STOICHEIA_DATA/parser_data/runs/joint_docclean_f0_s0 --split test
```
**Macronization and scansion (Tables 6-7).** The macronization ablation uses the
macron-only runs (`mac_v2*`, six seeds per arm), the scansion ablation the joint runs
(`joint_docclean*` / `joint_randinit*`); the external macronizer comparison scores the
joint model on Norma's 1,916 test positions:
```bash
python -m meter.predict --model $STOICHEIA_DATA/runs/meter_joint_docclean/best.pt --norma
python -m meter.predict --model $STOICHEIA_DATA/runs/meter_mac_v2/best.pt --norma --norma-source git
```
`python -m analysis.paper_tables` collects finished evaluation logs into the table
layouts used in the paper.
## 6. Tests
```bash
pytest tests/ -q
```
CPU-only, seconds-scale. Covers normalization/packing/noising (pretraining), edit-script
lemma encoding + dataset construction (tagger), macron/scansion mark parsing (meter),
plus lightweight forward-pass smoke tests for the restoration and joint tagger/parser
pipelines (tiny randomly-initialized configs — these check the tensor plumbing survives
refactors, not model quality).
## Reproducing the 10-fold split itself
The split is not an input -- it is constructed by the 13-stage pipeline in
`data/split_pipeline/` (sentence segmentation, pristine-edition clustering via
exact/MinHash-LSH/shared-sentence evidence, zone assignment, per-fold exclusion
masks, documentary decontamination incl. blocklist, bronze back-translation
filters, verification, manifests). `data/fold_manifests/SPLIT_DESIGN.md`
documents the design; `fold_k_test_works.tsv.gz` lists every work in fold k's
test bucket (group id, kind, source, record/char counts, author+title), so the
"provably never seen" guarantee is checkable work-by-work without rebuilding
anything.
## Revision naming
The paper's reconstruction revisions v1 and v2 correspond, for historical
reasons, to configuration files named `finetune_whole_v3_*` and
`finetune_whole_v4_*` respectively (earlier internal iterations v1/v2 were
superseded before evaluation and are not part of the release).
## External pieces the harnesses expect
* Meter training data: `MACRONIZER_SRC` must point at a checkout of the
`Ericu950/Stoicheia-meter-silver` dataset (silver verse lines + scanner corpus);
the Norma benchmark itself comes from `Ericu950/norma`.
* The Ithaca baseline harness (`insc/eval/ithaca_baseline.py`) expects DeepMind's
Ithaca repository cloned at `ithaca_upstream/` and its released checkpoint;
the Aeneas harness (`insc/eval/ptp_baseline.py`) expects DeepMind's
`predictingthepast` repository and its Greek checkpoint. Both are public
third-party releases and are not vendored here.
|