160M Dolma 3 Mix experiment checkpoints

Five 160M GPT-NeoX training runs on the separate M/C/G source pools in Dolma3-16B. Clean, M500, C500, G500, and MCG500 are separate run directories under runs/. Each contains all existing step-*.pt checkpoints, including the final step-122071.pt after 16,000,090,112 tokens. Intermediate checkpoints are approximately 1B tokens apart, with additional scheduled steps.

Standalone final weights are under models/<run>/model.safetensors, with a matching config and tokenizer in each run directory. Load one with GPTNeoXForCausalLM.from_pretrained("294-288-pretrain-interaction/160M-exp", subfolder="models/Clean") and the tokenizer from the same subfolder.

Each .pt is a PyTorch training checkpoint containing model, optimizer, scheduler, rank RNG states, step, token count, and run/manifest/config hashes. The .pt files are not direct from_pretrained artifacts. For trusted checkpoints, load the model state dict into GPTNeoXForCausalLM(GPTNeoXConfig.from_json_file("config.json")). Use the included tokenizer/. The training/ directory contains the training and data code; manifest.json fixes the source mixture, event schedule, and seeds.

M500, C500, and G500 inject 500 events from one source each; MCG500 divides 500 events across M/C/G as 167/167/166. The model weights are experimental research artifacts.

Downloads last month
8
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 294-288-pretrain-interaction/160M-exp

Finetuned
(380)
this model