zero4 / CORPUS_RIGHTS.md
Jonathan Beckwith
Publish ZERO.4
d53f2cd verified
|
Raw
History Blame Contribute Delete
6.87 kB

ZERO.4 corpus rights and provenance

Review date: 2026-08-12

Artifact: docs/model.litq8

SHA-256: 44b32f2262be2754fd2eeaf16ed206bae32b4ce30d7f5541a1059cd21257ae50

Release decision

ZERO.4 may be published as model weights under CC BY-SA 4.0. The checked memorization gate passed for every protected third-party stream. The Hugging Face model repository must use huggingface/release-manifest.json as an allowlist. It must not contain training text, token streams, raw downloads, or evaluation datasets.

This is a conservative rights and provenance record, not legal advice or a guarantee that the same copyright rules apply in every country. Creative Commons notes both that AI-training law varies and that using the same CC license for a publicly shared model is the conservative way to follow a ShareAlike source condition.

Bound training lineage

ZERO.4 was initialized from immutable ZERO.3 and trained with immutable ZERO.1, ZERO.2, and ZERO.3 teachers. The teacher hashes are bound in teachers/registry.json and corpus/RIGHTS.json. Its replay mixture consisted of the foundation, Shakespeare, Blake, Crowley, KJV, and literary-channel streams. Its added faculty data was produced by the checked quantity-request generator. The promoted checkpoint is Q2.6 seed 2, update 500.

The recorded lineage contains no human chat export. The channel stream was generated only from the named literary inputs. Later Q2.7/Q2.8 research and post-training external evaluations are not training sources for ZERO.4.

Source-level assessment

Training slice Source status Release treatment
ZERO foundation Project-authored statements CC0 1.0
Shakespeare Project Gutenberg eBook 100, identified there as public domain in the USA Attribute source; do not include text in the model repo
Blake eBook 574 and eBook 45315, identified there as public domain in the USA Attribute sources; do not include text in the model repo
Crowley: Tannhäuser and Household Gods eBook 70261 and eBook 14040, identified there as public domain in the USA Attribute sources; jurisdiction review required for dataset redistribution
Crowley: Clouds without Water Wikisource revision 13649032; underlying work marked public domain in the USA, transcription contributions under CC BY-SA Preserve revision, history, license, and change notice
Crowley: Liber AL vel Legis Wikisource revision 15225259; same license layers Preserve revision, history, license, and change notice; source document is reported as unknown
King James Bible Project Gutenberg eBook 30, identified there as public domain in the USA; special Crown publication rights apply in the UK Never include the KJV text in the Hugging Face package
Literary channel Mechanically derived from the Shakespeare, Blake, and Crowley streams Follows those inputs; no human chat data
Quantity requests Project-generated typed records CC0 1.0 to the extent rights exist

Project Gutenberg's license policy explains that its US-public-domain text is unrestricted by US copyright when the Gutenberg license/trademark wrapper is removed, while users outside the USA must check local law. Those wrappers were removed here. This project does not claim Project Gutenberg endorsement.

The UK government's copyright-term notice records the KJV's special letters-patent regime. Excluding the KJV text from the model package avoids representing it as a globally unrestricted dataset.

Transformations and attribution

The literary inputs underwent wrapper removal, UTF-8/LF normalization, character-level ASCII normalization, removal of specified editorial noise, and whitespace normalization. Wikisource material was therefore modified. The permanent revision and contributor-history links above provide reasonable attribution for the collaborative transcription layer; downstream dataset publication would require a separate, attribution-preserving review.

Original download and transformation hashes are in corpus/SHA256SUMS and corpus/RIGHTS.json. Raw downloads and generated token streams are deliberately gitignored; their hashes remain so the exact inputs can be reconstructed and checked without placing them in the model release.

Model and code licensing

  • Trained model artifacts: CC BY-SA 4.0, to the extent controlled rights apply.
  • Project code and runtime: Apache 2.0.
  • Eligible first-party generated data: CC0 1.0.
  • Literary text and derived records: source-specific status above; neither the Apache nor model license relicenses them.

This split avoids implying that an Apache software license clears the corpus. The CC BY-SA model license is conservative overcompliance; it is not a legal conclusion that trained weights are necessarily an adaptation in every jurisdiction.

Memorization release gate

The 2026-08-12 evaluation used sixteen evenly stratified windows per bound stream, a 128-token source prompt, and a 64-token greedy continuation. No protected third-party stream reached the 32-token warning threshold: the maxima were 2 for Shakespeare, 16 for Blake, 4 for Crowley, 6 for the KJV, and 9 for the derived literary channel.

Eight of sixteen project-authored foundation probes reproduced all 64 continuation tokens. That stream is intentionally inspectable, dedicated under CC0, and is therefore recorded as an informational rather than third-party-rights blocker. The complete hash-bound, text-free result is in release/zero4-memorization-v1.json.

Safety and limitations

The corpus contains archaic language plus sexual, violent, coercive, discriminatory, religious, and drug-related material, especially in the Crowley sources. ZERO.4 may reproduce themes or short phrases from its small corpus. It is not suitable as an authority on religion, history, medicine, law, identity, or factual questions. The memorization evaluation is a useful release gate, not proof that no source expression can ever be reproduced.