open-reason-large / README.md
theworker02's picture
Polish large CPU model card (91,544,064 params, host CPU, not 1B)
05c7782 verified
|
Raw
History Blame Contribute Delete
1.97 kB
metadata
language:
  - en
license: apache-2.0
library_name: transformers
tags:
  - open-reason
  - causal-lm
  - cpu
datasets:
  - theworker02/open-reason
base_model: gpt2-scratch

Open Reason large (CPU)

This is a large GPT-2-style causal LM trained from scratch on the Open Reason SFT split. It is larger than theworker02/open-reason-medium (13,867,008 parameters) and is not a 1B model and is not theworker02/open-reason-1b.

Weights live on this Hub repo. They are not stored in the GitHub git tree.

Training facts

  • Parameters: 91,544,064
  • Architecture: GPT-2-style from scratch; n_layer=12, n_embd=768, n_head=12, vocab_size=8192, max_seq_len=256
  • Steps: 400 (batch size 2)
  • Final training loss: 5.7361 (next-token NLL on training batches; not a benchmark score)
  • SFT rows: 3,175 from data/release/all.jsonl
  • Dataset: theworker02/open-reason pipeline 1.4.0
  • License: Apache-2.0
  • Hardware: AMD Ryzen 9 9950X host CPU; torch 2.12.0+cpu; torch.cuda.is_available()=False; Docker not installed and not used; nvidia-smi not present. NVIDIA CUDA was not used. AMD GPU / ROCm / DirectML were not used.

Related

No Reddit sources.

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("theworker02/open-reason-large")
model = AutoModelForCausalLM.from_pretrained("theworker02/open-reason-large")