| > **Published copy.** This file is the release copy of `data/SOURCES.md` from |
| > [https://github.com/naeyn/de-htr-web](https://github.com/naeyn/de-htr-web). Paths to files that ship in this release have been rewritten to |
| > their locations here (`training-data/…`). Any other repository-relative path — |
| > `src/…`, `scripts/…`, `DECISIONS.md`, the raw corpora, the font `.ttf` files, and |
| > `data/german_wikipedia_lines.txt` — refers to the source repository, which is the |
| > only place those files exist. |
|
|
| # License register — `de-htr-web` |
|
|
| **Project:** `de-htr-web` |
| **Purpose:** Record every dataset, font source, code dependency, external tool, and baseline model used by the project, with license id, source URL, and access date. |
| **Date:** 2026-07-28 |
|
|
| --- |
|
|
| ## Tier 1 — primary training data |
|
|
| Licenses below are copied from `.metadata.license.id` on the live Zenodo API record (`GET https://zenodo.org/api/records/<id>`), accessed 2026-07-28. |
|
|
| | Dataset | Creators | Zenodo record URL | DOI | Content | License | Confirmed on record page | Access date | |
| |---------|----------|-------------------|-----|---------|---------|--------------------------|-------------| |
| | ScaDS.AI German Line- and Word-Level Handwriting Dataset (v1.0) | Burghardt, Thomas; Alzin, Ahmad | https://zenodo.org/records/18301532 | 10.5281/zenodo.18301532 | 775 page images; 5,843 line images; 37,406 word images (the record does not state a writer count) | `cc-by-4.0` | yes (API `.metadata.license.id` = `cc-by-4.0`) | 2026-07-28 | |
| | ScaDS.AI German Full-Page Handwriting Dataset | Nestler, Till; Burghardt, Thomas | https://zenodo.org/records/18283705 | 10.5281/zenodo.18283705 | 77 pages; **not used** in the v1 release (see status note) | `cc-by-4.0` | yes (API `.metadata.license.id` = `cc-by-4.0`) | 2026-07-28 | |
|
|
| **Required CC BY 4.0 attribution.** CC BY 4.0 §3(a)(1)(A)(i) requires retaining |
| identification of the creators. The attribution statements for these two records are: |
|
|
| > "ScaDS.AI German Line- and Word-Level Handwriting Dataset" (v1.0, published 2026-01-19) |
| > by Thomas Burghardt and Ahmad Alzin, Zenodo, doi:10.5281/zenodo.18301532, licensed under |
| > [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/). |
|
|
| > "ScaDS.AI German Full-Page Handwriting Dataset" (published 2026-01-17) by Till Nestler |
| > and Thomas Burghardt, Zenodo, doi:10.5281/zenodo.18283705, licensed under |
| > [CC BY 4.0](https://creativecommons.org/licenses/by/4.0/). |
|
|
| Creator names were read from `.metadata.creators[].name` on the live Zenodo API records. |
|
|
| **Status note — full-page dataset.** The 77-page full-page set was reserved as a |
| held-out set for the optional M5 full-page milestone. M5 was skipped (DECISIONS.md D12), so |
| this record contributed **no training, validation, or test data** to the v1 release. It was |
| downloaded only to verify that it shares no writer with the line set |
| (`fullpage_writer_overlap == []` in `training-data/split_manifest.json`). Its attribution is retained |
| here because the data was obtained and inspected. |
|
|
| **Coverage note (measured, not from the record description):** the 775 pages describe the full |
| collection. Line-level ground truth exists for all 775 (775 XML files), but line **crops** exist |
| for only **385** of those pages, so the other 5,976 ground-truth lines are unusable as shipped and |
| are recorded as `xml_lines_skipped_missing_image` in `training-data/split_manifest.json`. The frozen |
| page-disjoint split therefore covers 385 pages / 5,843 lines (309/38/38 pages → |
| 4,675/567/601 lines). The 77-page full-page set shares **no** page identifier with the line set |
| (verified: `fullpage_writer_overlap == []`), so it remains a clean held-out set under the same |
| proxy caveat below. |
|
|
| **Writer-identity note (release-relevant).** The corpus ships **no writer identifiers**. The |
| split treats one page as one writer (`writer_id_field: page_id` in the manifest), so the split is |
| verifiably *page*-disjoint and only assumed to be *writer*-disjoint. The only identifying |
| metadata in the collection (`metadata/Uni_Leipzig_01-export-251205.csv`: `gender`, |
| `yearofbirth`) is too coarse to test the assumption — the 385 split pages fall into 90 such |
| profiles, up to 12 pages per profile, and every profile present in test also occurs in train. If |
| one contributor wrote several pages, some test hands may also appear in training and the reported |
| CER would be optimistic. This is stated in the README and the model card rather than left |
| implicit. |
|
|
| --- |
|
|
| ## Tier 2 — excluded datasets |
|
|
| None of these are used in v1. |
|
|
| | Dataset | Record | Reason excluded | |
| |---------|--------|-----------------| |
| | HTR German Kurrent 19th c. | Zenodo https://zenodo.org/records/17252677 | **MIXED** license: only the DTA subset is CC BY 4.0; Bayerische Schriftkunde transcriptions are CC BY-NC-SA 4.0 (`NC` forbids commercial use, `SA` is share-alike). Excluded entirely for v1 (also historical, out of scope). | |
| | READ dataset Bozen | Zenodo https://zenodo.org/records/218236 | License not verified; research-only risk — excluded | |
| | read_dataset_german_konzilsprotokolle | Zenodo https://zenodo.org/records/215383 | Not verified — excluded | |
| | StABS Ratsbücher O10 / Urfehdenbuch X | Zenodo https://zenodo.org/records/5153263 | Not verified — excluded | |
| | HTR_Validation_Set_StAZH_RRB_German_Kurrent_XIX | (Staatsarchiv Zürich; no verified Zenodo/record URL) | Not verified — excluded | |
| | ubtue/Ground-Truth Senatsprotokolle | GitHub | Not verified — excluded | |
| | IAM Handwriting Database | IAM registration | Modern English + restrictive research-only registration — excluded (wrong language, wrong license) | |
|
|
| --- |
|
|
| ## Tier 3 — synthetic data |
|
|
| Synthetic-data inputs: |
|
|
| - **Fonts:** see `training-data/FONTS.md`. SIL OFL 1.1 only. |
| - **German Wikipedia text corpus:** 10,234 clean rendering lines in |
| `data/german_wikipedia_lines.txt`, fetched reproducibly from 355 German Wikipedia |
| page revisions. `data/german_wikipedia_sources.json` records every title, page ID, |
| revision ID, and canonical URL. License: CC BY-SA 4.0 |
| (https://creativecommons.org/licenses/by-sa/4.0/); SA applies to the source text. |
| **`data/german_wikipedia_lines.txt` is committed to this repository and is therefore |
| redistributed under CC BY-SA 4.0, not under the repository's Apache-2.0 license.** The |
| full attribution and ShareAlike notice sits beside it in |
| `data/german_wikipedia_lines.LICENSE.txt`; see also "Files in this repository not covered |
| by Apache-2.0" below. |
| - **Real-word-image composition** from the 37,406 Tier-1 word images: license follows CC BY 4.0 (same as the Tier-1 word images). |
|
|
| **Hard rule:** synthetic data never appears in validation or test. |
|
|
| --- |
|
|
| ## Code dependencies |
|
|
| Versions from `uv.lock` / `package-lock.json`. Licenses from package metadata and the PyPI JSON API, accessed 2026-07-28 for M0/M1 entries and 2026-07-29 for M2–M4 additions. Ambiguous strings are recorded as published — not normalized. |
|
|
| | Name | Version | License | Package URL | |
| |------|---------|---------|-------------| |
| | pillow | 12.3.0 | `MIT-CMU` (`info.license_expression`; `info.license` null, no License classifier) | https://pypi.org/project/pillow/ | |
| | numpy | 2.5.1 | `BSD-3-Clause AND 0BSD AND MIT AND Zlib AND CC0-1.0` (`info.license_expression`; `info.license` null) | https://pypi.org/project/numpy/ | |
| | lxml | 6.1.1 | `BSD-3-Clause` (`info.license`) | https://pypi.org/project/lxml/ | |
| | rapidfuzz | 3.14.5 | `MIT` (`info.license_expression`; `info.license` null) | https://pypi.org/project/rapidfuzz/ | |
| | tqdm | 4.70.0 | `MPL-2.0 AND MIT` (`info.license`) | https://pypi.org/project/tqdm/ | |
| | requests | 2.34.2 | `Apache-2.0` (`info.license`; classifier `License :: OSI Approved :: Apache Software License`) | https://pypi.org/project/requests/ | |
| | fonttools | 4.63.0 | `MIT` (`info.license`) | https://pypi.org/project/fonttools/ | |
| | torch | 2.13.0 (current lock; **M2 training ran on 2.11.0+cu128** from the CUDA index — see the note below the table) | `Apache-2.0 AND Apache-2.0 WITH LLVM-exception AND BSD-2-Clause AND BSD-3-Clause AND BSL-1.0 AND MIT` (`info.license_expression`; `info.license` null) | https://pypi.org/project/torch/ | |
| | torchvision | 0.28.0 | `BSD` (`info.license` as published; ambiguous, not normalized) | https://pypi.org/project/torchvision/ | |
| | transformers | 5.14.1 | `Apache 2.0 License` (`info.license` as published) | https://pypi.org/project/transformers/ | |
| | sentencepiece | 0.2.2 | `Apache-2.0` (`info.license_expression`; `info.license` null) | https://pypi.org/project/sentencepiece/ | |
| | mlx-vlm | 0.6.8 | `MIT` (`info.license`; classifier `License :: OSI Approved :: MIT License`) | https://pypi.org/project/mlx-vlm/ | |
| | protobuf | 7.35.1 | `3-Clause BSD License` (`info.license` as published) | https://pypi.org/project/protobuf/ | |
| | onnx | 1.22.0 | `Apache-2.0` (`info.license_expression`; `info.license` null) | https://pypi.org/project/onnx/ | |
| | onnxruntime | 1.28.0 | `MIT License` (`info.license` as published) | https://pypi.org/project/onnxruntime/ | |
| | onnxruntime-node | 1.27.0 | `MIT` (`package.json`) | https://www.npmjs.com/package/onnxruntime-node | |
| | onnxruntime-web | 1.27.0 | `MIT` (`package.json`) | https://www.npmjs.com/package/onnxruntime-web | |
| | puppeteer | 25.4.0 | `Apache-2.0` (`package.json`; release verifier only) | https://www.npmjs.com/package/puppeteer | |
| | adm-zip | 0.6.0 | `MIT` (`package-lock.json`) — transitive dependency of `onnxruntime-node`, pinned forward from its `^0.5.16` range by the `overrides` block in `package.json` (DECISIONS.md D14); build/eval only, never shipped | https://www.npmjs.com/package/adm-zip | |
|
|
| **Versions differ between the lock files and the training runs.** The table records what |
| `uv.lock` / `package-lock.json` resolve today. M2 training ran on a rented Linux GPU host with |
| `torch 2.11.0+cu128` (CUDA 12.8) and `transformers 5.14.1`; the M3 PyTorch CPU control ran on |
| `torch 2.13.0`. Every run's exact environment is recorded in its own `results/**.json` under |
| `environment` / `config.environment`, which are the authoritative record for reproduction. |
|
|
| --- |
|
|
| ## External tools and models |
|
|
| | Name | Role | License | Source / API | Access date | |
| |------|------|---------|--------------|-------------| |
| | tesseract OCR 5.3.4 | Conventional OCR baseline (M1) | Apache-2.0 | https://github.com/tesseract-ocr/tesseract | 2026-07-28 | |
| | tessdata_best `deu` traineddata | German language data for Tesseract | Apache-2.0 | https://github.com/tesseract-ocr/tessdata_best | 2026-07-28 | |
| | **torchvision `ResNet18_Weights.IMAGENET1K_V1`** | **Pretrained visual stem of the released model — these weights are redistributed inside `models/dehtr_ctc_int8.onnx`** | Weights are published as part of the torchvision distribution under `BSD-3-Clause` (https://github.com/pytorch/vision/blob/main/LICENSE). See the ImageNet provenance note below. | https://github.com/pytorch/vision — used via `torchvision.models.resnet18(weights=ResNet18_Weights.DEFAULT)` (`src/dehtr/model.py`) | 2026-07-29 | |
| | microsoft/trocr-base-handwritten | TrOCR English handwritten zero-shot baseline (M1) | `mit` (`.cardData.license`) | https://huggingface.co/api/models/microsoft/trocr-base-handwritten | 2026-07-28 | |
| | microsoft/trocr-small-handwritten | Fine-tuned encoder-decoder **accuracy reference** (M2); **not released** | **UNVERIFIED / no declared license.** The HF API returns no `.cardData.license`, no `.cardData.license_name`, and no license tag. Upstream TrOCR code and checkpoints originate from `microsoft/unilm`, which is `MIT` (https://github.com/microsoft/unilm/blob/master/LICENSE), but the model repo itself asserts nothing. Recorded as UNVERIFIED rather than assumed. Used only to train a non-released comparison reference; no weights derived from it are distributed. | https://huggingface.co/api/models/microsoft/trocr-small-handwritten | 2026-07-29 | |
| | mlx-community/Qwen2.5-VL-7B-Instruct-4bit | VLM zero-shot baseline, MLX 4-bit (M1) | `apache-2.0` (`.cardData.license`) | https://huggingface.co/api/models/mlx-community/Qwen2.5-VL-7B-Instruct-4bit | 2026-07-28 | |
| | Qwen/Qwen2.5-VL-7B-Instruct | Base VLM for the MLX quantized variant | `apache-2.0` (`.cardData.license`) | https://huggingface.co/api/models/Qwen/Qwen2.5-VL-7B-Instruct | 2026-07-28 | |
| | ~~Qwen/Qwen2.5-VL-3B-Instruct~~ (rejected) | Considered as the VLM baseline, **not used** | **Non-permissive** — `.cardData.license` absent; `.cardData.license_name` = `qwen-research` (research-only). Rejected in favour of the apache-2.0 7B variant; see DECISIONS.md D2 | https://huggingface.co/api/models/Qwen/Qwen2.5-VL-3B-Instruct | 2026-07-28 | |
|
|
| ### ImageNet provenance note (release-relevant) |
|
|
| The released model's visual stem is initialized from torchvision's ImageNet-1k pretrained |
| ResNet-18 checkpoint, so those pretrained weights are **carried inside the released |
| artifact** — unlike the TrOCR and Qwen baselines, which are only evaluated against and never |
| redistributed. Two distinct licensing layers apply and are recorded separately: |
|
|
| 1. **The checkpoint file** is distributed by the torchvision project under torchvision's |
| `BSD-3-Clause` license. BSD-3-Clause permits commercial redistribution with attribution, |
| which is the basis on which this project's Apache-2.0 release includes those weights. |
| 2. **The ImageNet-1k images** the checkpoint was trained on are governed by the ImageNet |
| terms of access, which restrict the *images* to non-commercial research use. ImageNet |
| images are not redistributed by this project. |
|
|
| Whether restrictions on training images propagate to model weights trained on them is |
| legally unsettled and untested in court. The prevailing practice — and torchvision's own |
| distribution of these weights under BSD-3-Clause — treats the weights as freely reusable, |
| and this release relies on that reading. This is a **weaker guarantee** than the rest of the |
| register, where every input was chosen to be unambiguously permissive (DECISIONS.md D2). It is recorded explicitly here rather than left implicit. Anyone with a |
| stricter risk posture can retrain the stem from scratch: `CTCRecognizer(pretrained=False)` |
| in `src/dehtr/model.py` removes the ImageNet dependency entirely, at an accuracy cost that |
| this project has not measured. |
|
|
| --- |
|
|
| ## Third-party files redistributed in this repository |
|
|
| The Apache License 2.0 in `LICENSE` covers this project's own code and released model |
| artifacts, and its copyright line (`NOTICE`) speaks only for this project. The files below are |
| third-party works committed alongside it; they keep their own licenses and copyright holders. |
| Paths are relative to the GitHub source repository, and none of these files is part of the |
| Hugging Face model release: |
|
|
| | Path | License | Copyright holder | Notice | |
| |------|---------|------------------|--------| |
| | `data/german_wikipedia_lines.txt` | `CC BY-SA 4.0` | the respective German Wikipedia contributors | `data/german_wikipedia_lines.LICENSE.txt`, `data/german_wikipedia_sources.json` | |
| | `data/fonts/*.ttf` | `SIL OFL 1.1` | the respective font authors | `training-data/font-licenses/`, `training-data/FONTS.md` | |
| | `data/tessdata/deu.traineddata` | `Apache-2.0` (same license as this project, but **not** this project's copyright) | Google Inc. and the tesseract-ocr contributors | `NOTICE`; upstream https://github.com/tesseract-ocr/tessdata_best | |
| |
| `data/tessdata/deu.traineddata` is the unmodified `deu` model from `tessdata_best`, committed so |
| the M1 Tesseract floor baseline (`scripts/baseline_tesseract.py`) is reproducible without a |
| second download. It is used only as a baseline input: nothing derived from it enters the released |
| model. |
|
|
| Third-party assets redistributed by the **browser distribution** (ONNX Runtime Web, and the |
| Patrick Hand demo font) are recorded separately in `THIRD_PARTY_NOTICES.md`, with their |
| complete license texts under `third_party/`. |
|
|
| --- |
|
|
| Every entry is verified with a source URL and access date as noted above. Unverifiable items are marked UNVERIFIED rather than guessed. |
|
|