File size: 6,872 Bytes
d53f2cd
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
# ZERO.4 corpus rights and provenance

**Review date:** 2026-08-12

**Artifact:** `docs/model.litq8`

**SHA-256:** `44b32f2262be2754fd2eeaf16ed206bae32b4ce30d7f5541a1059cd21257ae50`

## Release decision

ZERO.4 may be published as model weights under **CC BY-SA 4.0**. The checked
memorization gate passed for every protected third-party stream. The Hugging
Face model repository must use `huggingface/release-manifest.json` as an
allowlist. It must not contain training text, token streams, raw downloads, or
evaluation datasets.

This is a conservative rights and provenance record, not legal advice or a
guarantee that the same copyright rules apply in every country. Creative
Commons notes both that AI-training law varies and that using the same CC
license for a publicly shared model is the conservative way to follow a
ShareAlike source condition.

## Bound training lineage

ZERO.4 was initialized from immutable ZERO.3 and trained with immutable
ZERO.1, ZERO.2, and ZERO.3 teachers. The teacher hashes are bound in
`teachers/registry.json` and `corpus/RIGHTS.json`. Its replay mixture consisted
of the foundation, Shakespeare, Blake, Crowley, KJV, and literary-channel
streams. Its added faculty data was produced by the checked quantity-request
generator. The promoted checkpoint is Q2.6 seed 2, update 500.

The recorded lineage contains **no human chat export**. The channel stream was
generated only from the named literary inputs. Later Q2.7/Q2.8 research and
post-training external evaluations are not training sources for ZERO.4.

## Source-level assessment

| Training slice | Source status | Release treatment |
| --- | --- | --- |
| ZERO foundation | Project-authored statements | CC0 1.0 |
| Shakespeare | [Project Gutenberg eBook 100](https://www.gutenberg.org/ebooks/100), identified there as public domain in the USA | Attribute source; do not include text in the model repo |
| Blake | [eBook 574](https://www.gutenberg.org/ebooks/574) and [eBook 45315](https://www.gutenberg.org/ebooks/45315), identified there as public domain in the USA | Attribute sources; do not include text in the model repo |
| Crowley: *Tannhäuser* and *Household Gods* | [eBook 70261](https://www.gutenberg.org/ebooks/70261) and [eBook 14040](https://www.gutenberg.org/ebooks/14040), identified there as public domain in the USA | Attribute sources; jurisdiction review required for dataset redistribution |
| Crowley: *Clouds without Water* | [Wikisource revision 13649032](https://en.wikisource.org/w/index.php?title=Clouds_without_Water&oldid=13649032); underlying work marked public domain in the USA, transcription contributions under CC BY-SA | Preserve revision, [history](https://en.wikisource.org/w/index.php?title=Clouds_without_Water&action=history), license, and change notice |
| Crowley: *Liber AL vel Legis* | [Wikisource revision 15225259](https://en.wikisource.org/w/index.php?title=Liber_AL_vel_Legis&oldid=15225259); same license layers | Preserve revision, [history](https://en.wikisource.org/w/index.php?title=Liber_AL_vel_Legis&action=history), license, and change notice; source document is reported as unknown |
| King James Bible | [Project Gutenberg eBook 30](https://www.gutenberg.org/ebooks/30), identified there as public domain in the USA; special Crown publication rights apply in the UK | Never include the KJV text in the Hugging Face package |
| Literary channel | Mechanically derived from the Shakespeare, Blake, and Crowley streams | Follows those inputs; no human chat data |
| Quantity requests | Project-generated typed records | CC0 1.0 to the extent rights exist |

Project Gutenberg's [license policy](https://www.gutenberg.org/policy/license)
explains that its US-public-domain text is unrestricted by US copyright when
the Gutenberg license/trademark wrapper is removed, while users outside the
USA must check local law. Those wrappers were removed here. This project does
not claim Project Gutenberg endorsement.

The UK government's [copyright-term notice](https://www.gov.uk/government/publications/copyright-notice-duration-of-copyright-term/copyright-notice-duration-of-copyright-term)
records the KJV's special letters-patent regime. Excluding the KJV text from
the model package avoids representing it as a globally unrestricted dataset.

## Transformations and attribution

The literary inputs underwent wrapper removal, UTF-8/LF normalization,
character-level ASCII normalization, removal of specified editorial noise,
and whitespace normalization. Wikisource material was therefore modified.
The permanent revision and contributor-history links above provide reasonable
attribution for the collaborative transcription layer; downstream dataset
publication would require a separate, attribution-preserving review.

Original download and transformation hashes are in `corpus/SHA256SUMS` and
`corpus/RIGHTS.json`. Raw downloads and generated token streams are deliberately
gitignored; their hashes remain so the exact inputs can be reconstructed and
checked without placing them in the model release.

## Model and code licensing

- Trained model artifacts: [CC BY-SA 4.0](LICENSE-MODEL.md), to the extent
  controlled rights apply.
- Project code and runtime: [Apache 2.0](https://www.apache.org/licenses/LICENSE-2.0).
- Eligible first-party generated data: [CC0 1.0](LICENSE-DATA.md).
- Literary text and derived records: source-specific status above; neither the
  Apache nor model license relicenses them.

This split avoids implying that an Apache software license clears the corpus.
The CC BY-SA model license is conservative overcompliance; it is not a legal
conclusion that trained weights are necessarily an adaptation in every
jurisdiction.

## Memorization release gate

The 2026-08-12 evaluation used sixteen evenly stratified windows per bound
stream, a 128-token source prompt, and a 64-token greedy continuation. No
protected third-party stream reached the 32-token warning threshold: the
maxima were 2 for Shakespeare, 16 for Blake, 4 for Crowley, 6 for the KJV, and
9 for the derived literary channel.

Eight of sixteen project-authored foundation probes reproduced all 64
continuation tokens. That stream is intentionally inspectable, dedicated under
CC0, and is therefore recorded as an informational rather than
third-party-rights blocker. The complete hash-bound, text-free result is in
`release/zero4-memorization-v1.json`.

## Safety and limitations

The corpus contains archaic language plus sexual, violent, coercive,
discriminatory, religious, and drug-related material, especially in the
Crowley sources. ZERO.4 may reproduce themes or short phrases from its small
corpus. It is not suitable as an authority on religion, history, medicine,
law, identity, or factual questions. The memorization evaluation is a useful
release gate, not proof that no source expression can ever be reproduced.