NanoGPT-OOM

Final submission for the CSE 251B NanoGPT contest.

  • Public validation perplexity: 26.9185
  • Public validation loss: 3.292814
  • Parameters: 98,939,904
  • Training data: FineWeb-Edu sample-10BT, mixed public text data, and a conservative low-LR continuation from the best checkpoint
  • Tokenizer/vocab: GPT-2 BPE, 50257 output logits

Repository contents:

  • checkpoint.pt: trained model checkpoint
  • model.py: model definition with the required load_model(checkpoint_path, device) interface
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support