How to use from
Docker Model Runner
docker model run hf.co/theworker02/open-reason-large
Quick Links

Open Reason large (CPU)

This is a large GPT-2-style causal LM trained from scratch on the Open Reason SFT split. It is larger than theworker02/open-reason-medium (13,867,008 parameters) and is not a 1B model and is not theworker02/open-reason-1b.

Weights live on this Hub repo. They are not stored in the GitHub git tree.

Training facts

  • Parameters: 91,544,064
  • Architecture: GPT-2-style from scratch; n_layer=12, n_embd=768, n_head=12, vocab_size=8192, max_seq_len=256
  • Steps: 400 (batch size 2)
  • Final training loss: 5.7361 (next-token NLL on training batches; not a benchmark score)
  • SFT rows: 3,175 from data/release/all.jsonl
  • Dataset: theworker02/open-reason pipeline 1.4.0
  • License: Apache-2.0
  • Hardware: AMD Ryzen 9 9950X host CPU; torch 2.12.0+cpu; torch.cuda.is_available()=False; Docker not installed and not used; nvidia-smi not present. NVIDIA CUDA was not used. AMD GPU / ROCm / DirectML were not used.

Related

No Reddit sources.

from transformers import AutoModelForCausalLM, AutoTokenizer
tok = AutoTokenizer.from_pretrained("theworker02/open-reason-large")
model = AutoModelForCausalLM.from_pretrained("theworker02/open-reason-large")
Downloads last month
174
Safetensors
Model size
91.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Dataset used to train theworker02/open-reason-large