|
Download README.md from Ne30Charm/tiny-ced: direct link, hf CLI and curl.
- Browser
- Download file 2.54 kB
-
https://huggingface.co/Ne30Charm/tiny-ced/resolve/main/README.md
- Command line
-
hf download hf://Ne30Charm/tiny-ced/README.md
-
curl -L -o README.md https://huggingface.co/Ne30Charm/tiny-ced/resolve/main/README.md
2.54 kB
| library_name: pytorch | |
| pipeline_tag: text-generation | |
| language: | |
| - en | |
| datasets: | |
| - roneneldan/TinyStories | |
| tags: | |
| - safetensors | |
| - custom-code | |
| # Tiny CED | |
| Tiny CED is a small custom PyTorch encoder-decoder language model trained from | |
| scratch on the English [TinyStories](https://huggingface.co/datasets/roneneldan/TinyStories) | |
| dataset. This package contains inference-only weights in Safetensors format. | |
| This is not a Transformers `PreTrainedModel`. Load it with the included | |
| `model.py` and `safetensors.torch.load_model` as shown below. | |
| ## Model details | |
| | Item | Value | | |
| | --- | ---: | | |
| | Independent parameters | 19,667,712 | | |
| | Vocabulary size | 8,192 | | |
| | Hidden size | 384 | | |
| | Encoder layers | 4 | | |
| | Decoder layers | 4 | | |
| | Attention heads | 6 | | |
| | Feed-forward size | 1,024 | | |
| | Local attention window | 64 | | |
| | Trained context length | 512 | | |
| | Weight dtype | FP32 | | |
| The token embedding and output head are tied. `model.safetensors` was exported | |
| with `safetensors.torch.save_model` so the shared tensor is stored only once. | |
| ## Usage | |
| Install the runtime dependencies: | |
| ```bash | |
| pip install -r requirements.txt | |
| ``` | |
| Run generation on CPU: | |
| ```bash | |
| python generate.py \ | |
| --device cpu \ | |
| --text "Once upon a time, a little rabbit" \ | |
| --max-new-tokens 100 | |
| ``` | |
| Use `--device cuda` when CUDA is available. Set `--temperature 0` for greedy | |
| decoding. | |
| ## Training and evaluation | |
| The exported weights come from the existing `best.pt` checkpoint; no retraining | |
| was performed during export. | |
| - Optimizer: AdamW | |
| - Peak learning rate: 3e-4 | |
| - Best-checkpoint training tokens: 40,004,782 | |
| - Best-checkpoint step: 5,856 | |
| - Validation loss: 1.896527 | |
| - Validation perplexity: 6.662715 | |
| Validation used the TinyStories validation split and the same 8,192-token BPE | |
| tokenizer included in this repository. | |
| ## Files | |
| - `model.safetensors`: inference weights | |
| - `model.py`: exact custom PyTorch architecture | |
| - `config.json`: architecture and token IDs | |
| - `tokenizer.json`: BPE tokenizer | |
| - `generate.py`: minimal generation example | |
| ## Limitations | |
| This is a small research model trained only on synthetic English children's | |
| stories. It may generate incorrect, repetitive, biased, unsafe, or otherwise | |
| inappropriate text. It is not suitable for factual, safety-critical, or | |
| production use without further evaluation. | |
| The source project did not specify a license for its code or weights. Choose an | |
| appropriate license and confirm the training-data terms before publishing this | |
| package publicly. TinyStories is listed as CDLA-Sharing-1.0 on its dataset page. | |