open-reason-xl / README.md
theworker02's picture
Honest XL model card: 443719680 params, 120 CPU steps, NLL 5.8116
df972ca verified
|
Raw
History Blame Contribute Delete
1.79 kB
metadata
language:
  - en
license: apache-2.0
library_name: transformers
tags:
  - open-reason
  - causal-lm
  - cpu
datasets:
  - theworker02/open-reason
base_model: gpt2-scratch

Open Reason XL (CPU)

GPT-2-style causal LM trained from scratch on the Open Reason SFT split. Exact parameter count: 443,719,680. This is not a 1B model and is not theworker02/open-reason-1b.

Training facts

  • Parameters: 443,719,680 (sum(p.numel() for p in model.parameters()))
  • Architecture: n_layer=22, n_embd=1280, n_head=20, vocab=8192, seq=256
  • Steps: 120 (batch 1, gradient accumulation 2, gradient checkpointing)
  • Device: host CPU (AMD Ryzen 9 9950X, 32 threads)
  • torch: 2.12.0+cpu; cuda_available=False; Docker not installed
  • NVIDIA CUDA was not used. AMD GPU / ROCm / DirectML were not used
  • Dataset: theworker02/open-reason pipeline 1.4.0
  • SFT rows: 3175 (data/release/all.jsonl)
  • Final training NLL: 5.8116
  • License: Apache-2.0
  • No held-out exact-match / coding / math benchmark scores are claimed

Related checkpoints (none of these is a 1B model):

No Reddit sources. Project license Apache-2.0.

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("theworker02/open-reason-xl")
model = AutoModelForCausalLM.from_pretrained("theworker02/open-reason-xl")