|
Download README.md from SlayerLab/Slayer149-balanced: direct link, hf CLI and curl.
- Browser
- Download file 1.9 kB
-
https://huggingface.co/SlayerLab/Slayer149-balanced/resolve/main/README.md
- Command line
-
hf download hf://SlayerLab/Slayer149-balanced/README.md
-
curl -L -o README.md https://huggingface.co/SlayerLab/Slayer149-balanced/resolve/main/README.md
1.9 kB
| language: en | |
| license: apache-2.0 | |
| library_name: pytorch | |
| pipeline_tag: text-generation | |
| tags: | |
| - experimental | |
| - from-scratch | |
| # Slayer149 Balanced | |
| Experimental 149,333,081-parameter English causal model trained **from scratch**. | |
| 26 layers, hidden width 576, FFN 2496, 9 query heads / 3 KV heads, tied 24,576-token | |
| byte-level BPE vocabulary, QK normalization and value residuals. Architecture | |
| informed by Qwen3 and AltSlate's Apache-2.0 Jugnu recipe; no Jugnu weights reused. | |
| This is a separate experiment from SlayerLab/Slayer149, which remains available. | |
| Training may still be in progress. See `reports/status.json`, training metrics and | |
| checkpoint manifests. No leaderboard rank or winning result is claimed. | |
| The final full GLINT-derived evaluation is `reports/evaluation.json`, when present. | |
| Competitor scores may use different evaluation conventions. | |
| `tokenizer.json` preserves Unicode/whitespace and uses EOS ID 0. Tokenizer selection | |
| and dataset provenance are recorded in reports. Training sources are pinned | |
| FineWeb-Edu, DCLM-edu, Cosmopedia-v2 and FineMath, with document-disjoint validation | |
| and exact-overlap filtering against the benchmark suite. Filtering is not a | |
| semantic contamination guarantee. No training text or credentials are uploaded. | |
| `checkpoints/<step>/training-state.pt` contains model and both optimizer states. | |
| Restore with the matching tokenizer, config, source files and training code. | |
| Final safetensors appear under the final checkpoint and at repository root. | |
| Use the supplied custom PyTorch loader; this is not yet an AutoModel package. | |
| ```python | |
| from load_model import load_model | |
| model, tokenizer = load_model("/path/to/downloaded/repository", device="cuda") | |
| ``` | |
| Goal: have a durable checkpoint and evaluation report by 15:00 Warsaw time on | |
| 2026-10-01, within the user's confirmed free GPU reservation. This is not a | |
| promise of model quality or benchmark rank. | |