GPT-15M - trained from scratch on 300M tokens
12.4M-param GPT-2-style model (6 layers, 384 hidden, 6 heads, tied embeddings, vocab 4095). Trained for 1 epoch on 300M tokens: 200M TinyStories + 100M ultrachat dialogues. Final held-out validation loss: 1.75 (TinyStories-valid).
Use with Transformers.js in the browser (ONNX) or PyTorch (model.pt).
- Downloads last month
- -