Slayer149 Balanced

Experimental 149,333,081-parameter English causal model trained from scratch. 26 layers, hidden width 576, FFN 2496, 9 query heads / 3 KV heads, tied 24,576-token byte-level BPE vocabulary, QK normalization and value residuals. Architecture informed by Qwen3 and AltSlate's Apache-2.0 Jugnu recipe; no Jugnu weights reused. This is a separate experiment from SlayerLab/Slayer149, which remains available.

Training may still be in progress. See reports/status.json, training metrics and checkpoint manifests. No leaderboard rank or winning result is claimed. The final full GLINT-derived evaluation is reports/evaluation.json, when present. Competitor scores may use different evaluation conventions.

tokenizer.json preserves Unicode/whitespace and uses EOS ID 0. Tokenizer selection and dataset provenance are recorded in reports. Training sources are pinned FineWeb-Edu, DCLM-edu, Cosmopedia-v2 and FineMath, with document-disjoint validation and exact-overlap filtering against the benchmark suite. Filtering is not a semantic contamination guarantee. No training text or credentials are uploaded.

checkpoints/<step>/training-state.pt contains model and both optimizer states. Restore with the matching tokenizer, config, source files and training code. Final safetensors appear under the final checkpoint and at repository root. Use the supplied custom PyTorch loader; this is not yet an AutoModel package.

from load_model import load_model
model, tokenizer = load_model("/path/to/downloaded/repository", device="cuda")

Goal: have a durable checkpoint and evaluation report by 15:00 Warsaw time on 2026-10-01, within the user's confirmed free GPU reservation. This is not a promise of model quality or benchmark rank.

Downloads last month
679
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support