File size: 1,903 Bytes
fd0936c
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
---
language: en
license: apache-2.0
library_name: pytorch
pipeline_tag: text-generation
tags:
- experimental
- from-scratch
---
# Slayer149 Balanced

Experimental 149,333,081-parameter English causal model trained **from scratch**.
26 layers, hidden width 576, FFN 2496, 9 query heads / 3 KV heads, tied 24,576-token
byte-level BPE vocabulary, QK normalization and value residuals. Architecture
informed by Qwen3 and AltSlate's Apache-2.0 Jugnu recipe; no Jugnu weights reused.
This is a separate experiment from SlayerLab/Slayer149, which remains available.

Training may still be in progress. See `reports/status.json`, training metrics and
checkpoint manifests. No leaderboard rank or winning result is claimed.
The final full GLINT-derived evaluation is `reports/evaluation.json`, when present.
Competitor scores may use different evaluation conventions.

`tokenizer.json` preserves Unicode/whitespace and uses EOS ID 0. Tokenizer selection
and dataset provenance are recorded in reports. Training sources are pinned
FineWeb-Edu, DCLM-edu, Cosmopedia-v2 and FineMath, with document-disjoint validation
and exact-overlap filtering against the benchmark suite. Filtering is not a
semantic contamination guarantee. No training text or credentials are uploaded.

`checkpoints/<step>/training-state.pt` contains model and both optimizer states.
Restore with the matching tokenizer, config, source files and training code.
Final safetensors appear under the final checkpoint and at repository root.
Use the supplied custom PyTorch loader; this is not yet an AutoModel package.

```python
from load_model import load_model
model, tokenizer = load_model("/path/to/downloaded/repository", device="cuda")
```

Goal: have a durable checkpoint and evaluation report by 15:00 Warsaw time on
2026-10-01, within the user's confirmed free GPU reservation. This is not a
promise of model quality or benchmark rank.