File size: 9,228 Bytes
efdf5ed | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 | ---
language:
- de
license: cc-by-sa-4.0
library_name: pytorch
tags:
- german
- pytorch
- transformer
- base-model
- text-generation
- from-scratch
- custom-code
datasets:
- coral-nlp/german-commons
pipeline_tag: text-generation
---
# German Commons 350M, step 10,000
**Author:** Hamid Wakili
**Contact:** hamid@ideployed.com
**Source code:** [ideployed/sovereign-german-llm](https://github.com/ideployed/sovereign-german-llm)
This repository contains an intermediate checkpoint of my German base-model
experiment. I trained the tokenizer and model from random initialization. The
release does not use an imported tokenizer vocabulary, pretrained weights,
distillation, or synthetic training text.
The checkpoint was saved after 10,000 optimizer steps, equivalent to
5,242,880,000 training tokens at the configured global batch size. The prepared
corpus contains 30.00B training tokens, but this checkpoint has not completed
that training target.
This is a base model for text completion, not a chatbot. It has not been
instruction-tuned and should not be expected to answer questions or follow
commands reliably.
## Release contents
- `model.safetensors`: inference weights from step 10,000
- `config.json`: architecture configuration
- `german_transformer_model.py`: custom PyTorch model implementation
- `bpe_tokenizer.py`: custom byte-level BPE implementation
- `inference.py`: offline Safetensors generation CLI
- `tokenizer/de_bpe_32k/`: matching vocabulary, merges, and special tokens
- `samples/generations_5k_vs_10k.md`: deterministic checkpoint comparison
- `samples/training_curves/`: loss, perplexity, optimization, and throughput plots
- `SHA256SUMS`: integrity hashes for the model, runtime code, and tokenizer
The optimizer and scheduler states are intentionally excluded. The original
resumable checkpoint is a separate training artifact.
## Model architecture
| Property | Value |
|---|---:|
| Architecture | Decoder-only Transformer |
| Parameters | 341,885,952 |
| Layers | 24 |
| Model width | 1,024 |
| Attention | 16-head MHA, head dimension 64 |
| Feed-forward network | SwiGLU, width 2,816 |
| Normalization | Pre-RMSNorm |
| Positions | RoPE, theta 10,000 |
| Context length | 2,048 tokens |
| Vocabulary | 32,768-token byte-level BPE |
| Embeddings | Tied input/output embeddings |
| Dropout | 0.0 |
| Training precision | bfloat16 |
I implemented the architecture directly with PyTorch. It is not registered
with the Hugging Face Transformers `AutoModel` classes, so `trust_remote_code`
and `AutoModelForCausalLM` are not supported by this release. The complete
training and data-preparation code is available in the
[GitHub repository](https://github.com/ideployed/sovereign-german-llm).
## Tokenizer
The tokenizer is a byte-level BPE tokenizer trained for this experiment on a
sample from German Commons. Its fixed IDs are:
| ID | Token |
|---:|---|
| 0 | `<unk>` |
| 1 | `<pad>` |
| 2 | `<s>` |
| 3 | `</s>` |
| 4 | `</w>` |
IDs 5 through 260 represent the 256 byte values. Higher IDs represent learned
BPE merges. The tokenizer uses `</w>` as an explicit word boundary marker.
## Training data
The prepared data was streamed from
[German Commons](https://huggingface.co/datasets/coral-nlp/german-commons),
quality-filtered, domain-reweighted, and encoded into local `uint16` shards.
The prepared dataset contains:
- 30,000,000,243 training tokens
- 50,002,350 validation tokens
- a stable-hash document split with seed 1337
- aggregate source and license counts recorded during preparation
The accepted-document license distribution recorded in the preparation
manifest was approximately 51.6% CC0-1.0, 46.2% CC-BY-4.0, and 2.2%
CC-BY-SA-4.0. These percentages count accepted documents, not tokens.
German Commons contains a substantial amount of historical newspaper and
cultural material. Domain reweighting reduced that concentration but did not
remove it. The resulting model frequently uses historical spelling and register.
## Training configuration
| Property | Value |
|---|---:|
| Checkpoint step | 10,000 |
| Tokens consumed | 5.243B |
| Prepared training target | 30B tokens |
| Optimizer | AdamW |
| Adam betas | 0.9, 0.95 |
| Weight decay | 0.1 |
| Gradient clipping | 1.0 |
| Peak learning rate | 3e-4 |
| Warmup | 2,000 steps |
| Schedule | Cosine decay to 3e-5 |
| Global batch | 524,288 tokens per step |
| Micro-batch | 16 sequences |
| Gradient accumulation | 16 |
| Hardware | 1 NVIDIA H100 PCIe 80 GB |
| Measured throughput | approximately 60,000 tokens/second |
| Compile mode | disabled for this run |
At step 10,000, the recorded training loss was 3.2307. Evaluation over 20
batches from the held-out validation shards produced a validation loss of
3.2593 and perplexity of 26.03. This is a training diagnostic, not a broad
language-model benchmark.




## Evaluation status
No standardized downstream benchmark results are reported for this checkpoint.
The available evaluation consists of held-out German Commons loss and direct
inspection of deterministic completions. The 5k and 10k comparison is published
in [`samples/generations_5k_vs_10k.md`](samples/generations_5k_vs_10k.md).
Those samples are intentionally unedited. They show improved German structure
at step 10,000, but also repetition, historical register, and severe factual
fabrication.
## Usage
Install PyTorch and Safetensors, clone or download this repository, and run from
its root:
```bash
python inference.py \
--model-dir . \
--prompt "Die deutsche Sprache ist" \
--max-new-tokens 80 \
--temperature 0.8 \
--top-k 40 \
--top-p 0.9
```
The underlying loading path is plain PyTorch:
```python
import json
import torch
from safetensors.torch import load_file
from bpe_tokenizer import BPE_Tokenizer
from german_transformer_model import GermanGPT, GermanGPTConfig
device = torch.device("cuda" if torch.cuda.is_available() else "cpu")
with open("config.json", "r", encoding="utf-8") as file:
config = GermanGPTConfig.from_dict(json.load(file))
tokenizer = BPE_Tokenizer.load("tokenizer/de_bpe_32k")
model = GermanGPT(config)
state_dict = load_file("model.safetensors", device="cpu")
model.load_state_dict(state_dict, strict=True)
model.to(device).eval()
prompt = "Die deutsche Sprache ist"
input_ids = torch.tensor([tokenizer.encode(prompt)], dtype=torch.long, device=device)
with torch.inference_mode():
logits, _ = model(input_ids)
next_token = int(torch.argmax(logits[0, -1]).item())
print(tokenizer.decode(input_ids[0].tolist() + [next_token]))
```
This example performs one greedy next-token step. Autoregressive sampling can
be implemented by repeatedly passing the most recent 2,048 tokens through the
model and applying temperature, top-k, or top-p filtering.
## Intended use
- Research on German base-model pretraining and tokenization
- Inspection of training behavior at an early checkpoint
- Continuation training and controlled fine-tuning experiments
- Non-factual German text-completion experiments
## Limitations
1. **Intermediate checkpoint.** The model has consumed about 17.5% of the
prepared 30B-token target.
2. **Not instruction-tuned.** Question answering and chat prompts are outside
the model's training objective.
3. **Unreliable facts.** The model confidently invents people, places, dates,
quantities, and relationships. Its output must not be treated as evidence.
4. **Repetition.** Longer generations often enter repeated phrases or sentence
structures.
5. **Historical register.** Archaic spelling and phrasing occur frequently due
to the composition of the accepted training documents.
6. **German-focused.** Other languages were not evaluated and are outside the
intended scope.
7. **Inherited corpus bias.** Filtering and provenance records do not remove
social, historical, geographic, or source-selection bias.
8. **No safety evaluation.** I have not completed red-team, memorization,
privacy, toxicity, or demographic-bias evaluations.
Do not use this checkpoint for medical, legal, financial, emergency, or other
high-impact decisions. It is not suitable for autonomous decision-making about
people.
## Licensing and attribution
- The project code in this repository is licensed under the MIT License; see
`LICENSE`.
- The model weights are released under CC BY-SA 4.0 as stated in the repository
metadata.
- German Commons and each constituent source retain their own attribution and
licensing requirements. Users are responsible for reviewing those terms for
their intended use.
Please attribute both Hamid Wakili for this model release and German Commons as
the training-data source.
## Citation
```bibtex
@misc{wakili2026germancommons350m,
author = {Hamid Wakili},
title = {German Commons 350M, Step 10,000},
year = {2026},
note = {Intermediate German base-model checkpoint},
url = {https://github.com/ideployed/sovereign-german-llm}
}
```
|