Upload folder using huggingface_hub
Browse files- README.md +104 -0
- SHA256SUMS.txt +1 -0
- arch_facts.json +29 -0
- config.json +10 -0
- model.safetensors +3 -0
- tokenizer.json +0 -0
README.md
ADDED
|
@@ -0,0 +1,104 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
---
|
| 2 |
+
license: mit
|
| 3 |
+
language:
|
| 4 |
+
- code
|
| 5 |
+
tags:
|
| 6 |
+
- code-generation
|
| 7 |
+
- python
|
| 8 |
+
- from-scratch
|
| 9 |
+
- pretrained
|
| 10 |
+
library_name: pytorch
|
| 11 |
+
---
|
| 12 |
+
|
| 13 |
+
# code-llm-435m β a Python code-completion model trained from scratch on one RTX 5090
|
| 14 |
+
|
| 15 |
+
A 435M-parameter decoder-only Transformer for Python code completion. Data pipeline, tokenizer,
|
| 16 |
+
architecture, training loop, checkpoint merging and evaluation were all written and run by one person
|
| 17 |
+
on one consumer GPU β this repository holds the weights; the code, design record, ablation report and
|
| 18 |
+
failure log live in the GitHub repository linked below.
|
| 19 |
+
|
| 20 |
+
> **This folder is the pretrained base.** Its weights are the *merged* (weight-averaged) product of the
|
| 21 |
+
> pretraining run: under this training scheme the merged checkpoint, not the last step, is the finished
|
| 22 |
+
> pretrained model β hence the folder name. The instruction-tuned variant is in [`../sft`](../sft).
|
| 23 |
+
|
| 24 |
+
## Architecture
|
| 25 |
+
|
| 26 |
+
| | |
|
| 27 |
+
|---|---|
|
| 28 |
+
| Parameters | **434,680,832** (bf16) |
|
| 29 |
+
| Layers | 22 |
|
| 30 |
+
| d_model / heads | 1024 / 16 (head_dim 64) |
|
| 31 |
+
| FFN | SwiGLU, d_ff 4096 |
|
| 32 |
+
| Position | RoPE, theta 500000 |
|
| 33 |
+
| Norm | RMSNorm (pre-norm) |
|
| 34 |
+
| Attention | causal SDPA (no biases anywhere) |
|
| 35 |
+
| Embeddings | untied (input embed + output head) |
|
| 36 |
+
| Vocabulary | 32,000 byte-level BPE, trained on the filtered corpus (β3.5 chars/token) |
|
| 37 |
+
| Context | 1024 tokens |
|
| 38 |
+
| Precision | bf16 (released file) |
|
| 39 |
+
|
| 40 |
+
## Training
|
| 41 |
+
|
| 42 |
+
- **Data**: three Python sources (codeparrot / the_stack / star_coder), six-layer quality filtering,
|
| 43 |
+
copyright filtering, metadata stripping; concatenated into **one continuous 31.0 B-token stream**
|
| 44 |
+
(no dataset boundaries, no optimizer resets β an earlier version was silently retrained on the same
|
| 45 |
+
8 B tokens twice by a resume bug, and the design change makes that class of bug impossible).
|
| 46 |
+
- **Schedule**: WSM β constant learning rate, then a 30,000-step cooldown, then **weighted averaging
|
| 47 |
+
of the last 10,000 checkpoints** (this file is that merged model, not the last step).
|
| 48 |
+
- **Hardware**: a single RTX 5090, 32 GB, Blackwell sm_120.
|
| 49 |
+
|
| 50 |
+
## What it can and cannot do
|
| 51 |
+
|
| 52 |
+
**Can**: complete short Python functions when given a signature and the opening indentation.
|
| 53 |
+
On a human-graded benchmark (10 docstring-free tasks Γ 5 seeds, scored 0/1/2, max 100) this model
|
| 54 |
+
scores **66/100**; the earlier 353M version β same architecture, same GPU β scores **6/100**. The
|
| 55 |
+
difference is the data pipeline, not the architecture. Sample, verbatim (temperature 0.2, seed 0;
|
| 56 |
+
this is a partial-credit example, not a showcase):
|
| 57 |
+
|
| 58 |
+
```python
|
| 59 |
+
def quicksort(arr):
|
| 60 |
+
if len(arr) <= 1:
|
| 61 |
+
return arr
|
| 62 |
+
else:
|
| 63 |
+
pivot = arr[0]
|
| 64 |
+
left = [x for x in arr if x < pivot]
|
| 65 |
+
right = [x for x in arr if x == pivot] # <- wrong: should be > pivot
|
| 66 |
+
return quicksort(left) + [pivot] + quicksort(right)
|
| 67 |
+
```
|
| 68 |
+
|
| 69 |
+
The recursion, the base case and the partition are there; one comparison operator is wrong, so the
|
| 70 |
+
function drops elements. That is what "66/100" looks like at this scale β the structure is learned
|
| 71 |
+
before the detail is, which is exactly why the benchmark is human-graded rather than pass/fail.
|
| 72 |
+
|
| 73 |
+
**Cannot**: follow instructions β this is a completion model, not a chat model (supervised
|
| 74 |
+
fine-tuning is documented separately in the GitHub repo). It is blind to docstrings: at this scale a
|
| 75 |
+
from-scratch model reads a docstring as the end of the function, so standard HumanEval
|
| 76 |
+
docstring prompts score β0 and docstring-free prompts are used instead. At 18 B tokens the same
|
| 77 |
+
pipeline still scored 0/50 on a five-algorithm suite β implementation ability appears between
|
| 78 |
+
18 B and 31 B tokens, which the ablation report documents rather than hides.
|
| 79 |
+
|
| 80 |
+
## Files
|
| 81 |
+
|
| 82 |
+
| File | What |
|
| 83 |
+
|---|---|
|
| 84 |
+
| `model.safetensors` | bf16 weights, 157 tensors β verified bit-identical to the training checkpoint after the fp32βbf16 cast |
|
| 85 |
+
| `config.json` | architecture config exactly as stored in the training checkpoint |
|
| 86 |
+
| `tokenizer.json` | byte-level BPE, 32,000 tokens |
|
| 87 |
+
| `SHA256SUMS.txt` | artifact hash |
|
| 88 |
+
|
| 89 |
+
Loading it requires the model class from the training repository (`src/train.py`, class `CodeLLM`
|
| 90 |
+
with `ModelConfig(**config.json)`); `torch.load` of a state dict built by hand will not reproduce the
|
| 91 |
+
forward pass described above.
|
| 92 |
+
|
| 93 |
+
## Intended use and limits
|
| 94 |
+
|
| 95 |
+
Research and education: understanding what a few-hundred-million-parameter model actually learns when
|
| 96 |
+
trained end to end on real data. Not for production code generation, not for instruction following,
|
| 97 |
+
not a substitute for a competent developer. Trained only on code; no personal data. Outputs may
|
| 98 |
+
reproduce patterns (and licensing quirks) from the training corpus despite copyright filtering.
|
| 99 |
+
|
| 100 |
+
## Links
|
| 101 |
+
|
| 102 |
+
Everything else β `DESIGN.md` (decisions and why), `CHALLENGES.md` (every bug and contaminant),
|
| 103 |
+
the ablation report, the fine-tuning log with corrections, and every evaluated generation β is in the
|
| 104 |
+
GitHub repository. MIT licensed.
|
SHA256SUMS.txt
ADDED
|
@@ -0,0 +1 @@
|
|
|
|
|
|
|
| 1 |
+
d2c030b0fbd8c9a6a227b96e704a8c538c3754f468b2ebd2d9970f880b8c7b0d model.safetensors
|
arch_facts.json
ADDED
|
@@ -0,0 +1,29 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"params_total": 434680832,
|
| 3 |
+
"n_tensors": 157,
|
| 4 |
+
"dtype": "bf16",
|
| 5 |
+
"vocab_size": null,
|
| 6 |
+
"d_model": null,
|
| 7 |
+
"n_blocks": 22,
|
| 8 |
+
"tensor_keys_sample": [
|
| 9 |
+
"blocks.0.attn.out_proj.weight",
|
| 10 |
+
"blocks.0.attn.qkv.weight",
|
| 11 |
+
"blocks.0.attn_norm.scale",
|
| 12 |
+
"blocks.0.ffn.down.weight",
|
| 13 |
+
"blocks.0.ffn.gate.weight",
|
| 14 |
+
"blocks.0.ffn.up.weight",
|
| 15 |
+
"blocks.0.ffn_norm.scale",
|
| 16 |
+
"blocks.1.attn.out_proj.weight"
|
| 17 |
+
],
|
| 18 |
+
"source_checkpoint": "the pretrained checkpoint this export was taken from (path omitted in the released copy)",
|
| 19 |
+
"source_keys": [
|
| 20 |
+
"optimizer",
|
| 21 |
+
"scheduler",
|
| 22 |
+
"config",
|
| 23 |
+
"training_state",
|
| 24 |
+
"timestamp",
|
| 25 |
+
"model"
|
| 26 |
+
],
|
| 27 |
+
"step": null,
|
| 28 |
+
"created_utc": "2026-09-20T18:14:35Z"
|
| 29 |
+
}
|
config.json
ADDED
|
@@ -0,0 +1,10 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
{
|
| 2 |
+
"vocab_size": 32000,
|
| 3 |
+
"d_model": 1024,
|
| 4 |
+
"num_layers": 22,
|
| 5 |
+
"num_heads": 16,
|
| 6 |
+
"d_ff": 4096,
|
| 7 |
+
"max_seq_len": 1024,
|
| 8 |
+
"rope_theta": 500000.0,
|
| 9 |
+
"dropout_rate": 0.0
|
| 10 |
+
}
|
model.safetensors
ADDED
|
@@ -0,0 +1,3 @@
|
|
|
|
|
|
|
|
|
|
|
|
|
| 1 |
+
version https://git-lfs.github.com/spec/v1
|
| 2 |
+
oid sha256:d2c030b0fbd8c9a6a227b96e704a8c538c3754f468b2ebd2d9970f880b8c7b0d
|
| 3 |
+
size 869377400
|
tokenizer.json
ADDED
|
The diff for this file is too large to render.
See raw diff
|
|
|