PyGPT — a 110M-param Python code model, trained from scratch

GPT-2 Small architecture (12L/12H/768d, ctx 512, 32k codeparrot BPE vocab) trained from random initialisation for a university ML course project: 2.35B tokens of Python (codeparrot-clean, ~Chinchilla-optimal for this size) with a textbook-algorithms curriculum blended at ~6.6% of batches, 17.5 h on Kaggle 2x T4. Not fine-tuned from any pretrained model.

Held-out Python perplexity 4.52 (val loss 1.509) · 60% of sampled completions parse as valid Python (vanilla GPT-2 124M: 7%) · pass@1 0.60 / pass@10 0.80 on 5 classic DSA tasks (GPT-2: 0.00) — add 30/30 and divide 29/30 correct at n=30; is_prime remains 0, the honest boundary at this scale.

Prompting tip: use type hints + a triple-quoted docstring with doctest examples, and do NOT end the prompt with a trailing newline — the BPE fuses newline+indent into one token, so a bare newline tells the model the function is finished.

Run the demo locally (CPU is fine, ~40 tok/s)

pip install torch numpy transformers gradio
python app.py          # Gradio UI at http://127.0.0.1:7860

Files: model_fp16.pt (fp16 weights, 211 MB) · tokenizer/ · model.py (architecture, KV-cache generation) · app.py (demo) · training/eval code at https://github.com/Rehan123-bash/pygpt

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support