Mini Code Generator V7.3
A small Python code-generation model trained entirely from scratch in PyTorch.
V7.3 maps a natural-language programming instruction to a complete Python function.
Model
- Parameters: 10,742,016
- Architecture: causal Transformer
- Layers: 6
- Hidden size: 384
- Attention heads: 6
- FFN size: 1536
- Context length: 96 tokens
- Dropout: 0
- Tied input/output embeddings
- Vocabulary: 150 tokens
- Training: from scratch
- Pretrained model: none
- Pretrained tokenizer: none
Dataset
The controlled V7.3 benchmark contains:
- 30 programming families
- 20 natural-language instructions per family
- 600 total instruction/code records
- 30 canonical Python implementations
- one canonical implementation per family
- random record split
- 480 train
- 60 validation
- 60 test
The dataset was generated and validated locally before training.
Training
Best checkpoint:
- best epoch: 27
- best validation loss: 0.005979
- final training accuracy: approximately 99.95%
- final validation accuracy: approximately 99.70%
Training used teacher-forced next-token prediction, while the final benchmark uses free autoregressive generation.
Benchmark Result
Held-out test set:
| Metric | Result |
|---|---|
| Test prompts | 60 |
| Syntax-valid generations | 57/60 |
| Exact matches | 56/60 |
| Fully correct prompts | 56/60 (93.33%) |
| Hidden tests | 300 |
| Hidden tests passed | 283/300 (94.33%) |
| Executed hidden tests | 284 |
| Executed-test accuracy | 283/284 (99.65%) |
A prompt is counted as fully correct only when the generated program passes all five hidden functional tests associated with that prompt.
Important Scope
This is a controlled small-scale benchmark, not a claim of general coding ability.
The benchmark contains short algorithmic Python functions and only 60 held-out test prompts. The result should therefore be interpreted as a strong result for this specific from-scratch experiment.
Known Failures
The four non-fully-correct prompts involved:
product_listfind_firstcount_wordsis_palindrome
The detailed generated outputs and hidden-test results are included in the repository.
Files
checkpoints/v7_3_best.ptβ best validation checkpointcheckpoints/v7_3_final.ptβ final checkpointtokenizer/v7_3_tokenizer.jsonβ custom tokenizeranalysis/v7_3_final_report.jsonβ final experiment reportanalysis/v7_3_final_report.txtβ human-readable reportevaluation/v7_3_greedy_generations.jsonβ generated test programsevaluation/v7_3_hidden_functional_results.jsonβ hidden-test resultslogs/v7_3_training_history.jsonβ training history
Reproducibility
The model was trained from randomly initialized weights with a fixed seed.
No pretrained language model or pretrained code model was used.
This repository represents the frozen V7.3 baseline before subsequent experiments.