TightX commited on
Commit
32e9d3c
Β·
verified Β·
1 Parent(s): 4a8de13

Upload folder using huggingface_hub

Browse files
Files changed (6) hide show
  1. README.md +104 -0
  2. SHA256SUMS.txt +1 -0
  3. arch_facts.json +29 -0
  4. config.json +10 -0
  5. model.safetensors +3 -0
  6. tokenizer.json +0 -0
README.md ADDED
@@ -0,0 +1,104 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ ---
2
+ license: mit
3
+ language:
4
+ - code
5
+ tags:
6
+ - code-generation
7
+ - python
8
+ - from-scratch
9
+ - pretrained
10
+ library_name: pytorch
11
+ ---
12
+
13
+ # code-llm-435m β€” a Python code-completion model trained from scratch on one RTX 5090
14
+
15
+ A 435M-parameter decoder-only Transformer for Python code completion. Data pipeline, tokenizer,
16
+ architecture, training loop, checkpoint merging and evaluation were all written and run by one person
17
+ on one consumer GPU β€” this repository holds the weights; the code, design record, ablation report and
18
+ failure log live in the GitHub repository linked below.
19
+
20
+ > **This folder is the pretrained base.** Its weights are the *merged* (weight-averaged) product of the
21
+ > pretraining run: under this training scheme the merged checkpoint, not the last step, is the finished
22
+ > pretrained model β€” hence the folder name. The instruction-tuned variant is in [`../sft`](../sft).
23
+
24
+ ## Architecture
25
+
26
+ | | |
27
+ |---|---|
28
+ | Parameters | **434,680,832** (bf16) |
29
+ | Layers | 22 |
30
+ | d_model / heads | 1024 / 16 (head_dim 64) |
31
+ | FFN | SwiGLU, d_ff 4096 |
32
+ | Position | RoPE, theta 500000 |
33
+ | Norm | RMSNorm (pre-norm) |
34
+ | Attention | causal SDPA (no biases anywhere) |
35
+ | Embeddings | untied (input embed + output head) |
36
+ | Vocabulary | 32,000 byte-level BPE, trained on the filtered corpus (β‰ˆ3.5 chars/token) |
37
+ | Context | 1024 tokens |
38
+ | Precision | bf16 (released file) |
39
+
40
+ ## Training
41
+
42
+ - **Data**: three Python sources (codeparrot / the_stack / star_coder), six-layer quality filtering,
43
+ copyright filtering, metadata stripping; concatenated into **one continuous 31.0 B-token stream**
44
+ (no dataset boundaries, no optimizer resets β€” an earlier version was silently retrained on the same
45
+ 8 B tokens twice by a resume bug, and the design change makes that class of bug impossible).
46
+ - **Schedule**: WSM β€” constant learning rate, then a 30,000-step cooldown, then **weighted averaging
47
+ of the last 10,000 checkpoints** (this file is that merged model, not the last step).
48
+ - **Hardware**: a single RTX 5090, 32 GB, Blackwell sm_120.
49
+
50
+ ## What it can and cannot do
51
+
52
+ **Can**: complete short Python functions when given a signature and the opening indentation.
53
+ On a human-graded benchmark (10 docstring-free tasks Γ— 5 seeds, scored 0/1/2, max 100) this model
54
+ scores **66/100**; the earlier 353M version β€” same architecture, same GPU β€” scores **6/100**. The
55
+ difference is the data pipeline, not the architecture. Sample, verbatim (temperature 0.2, seed 0;
56
+ this is a partial-credit example, not a showcase):
57
+
58
+ ```python
59
+ def quicksort(arr):
60
+ if len(arr) <= 1:
61
+ return arr
62
+ else:
63
+ pivot = arr[0]
64
+ left = [x for x in arr if x < pivot]
65
+ right = [x for x in arr if x == pivot] # <- wrong: should be > pivot
66
+ return quicksort(left) + [pivot] + quicksort(right)
67
+ ```
68
+
69
+ The recursion, the base case and the partition are there; one comparison operator is wrong, so the
70
+ function drops elements. That is what "66/100" looks like at this scale β€” the structure is learned
71
+ before the detail is, which is exactly why the benchmark is human-graded rather than pass/fail.
72
+
73
+ **Cannot**: follow instructions β€” this is a completion model, not a chat model (supervised
74
+ fine-tuning is documented separately in the GitHub repo). It is blind to docstrings: at this scale a
75
+ from-scratch model reads a docstring as the end of the function, so standard HumanEval
76
+ docstring prompts score β‰ˆ0 and docstring-free prompts are used instead. At 18 B tokens the same
77
+ pipeline still scored 0/50 on a five-algorithm suite β€” implementation ability appears between
78
+ 18 B and 31 B tokens, which the ablation report documents rather than hides.
79
+
80
+ ## Files
81
+
82
+ | File | What |
83
+ |---|---|
84
+ | `model.safetensors` | bf16 weights, 157 tensors β€” verified bit-identical to the training checkpoint after the fp32β†’bf16 cast |
85
+ | `config.json` | architecture config exactly as stored in the training checkpoint |
86
+ | `tokenizer.json` | byte-level BPE, 32,000 tokens |
87
+ | `SHA256SUMS.txt` | artifact hash |
88
+
89
+ Loading it requires the model class from the training repository (`src/train.py`, class `CodeLLM`
90
+ with `ModelConfig(**config.json)`); `torch.load` of a state dict built by hand will not reproduce the
91
+ forward pass described above.
92
+
93
+ ## Intended use and limits
94
+
95
+ Research and education: understanding what a few-hundred-million-parameter model actually learns when
96
+ trained end to end on real data. Not for production code generation, not for instruction following,
97
+ not a substitute for a competent developer. Trained only on code; no personal data. Outputs may
98
+ reproduce patterns (and licensing quirks) from the training corpus despite copyright filtering.
99
+
100
+ ## Links
101
+
102
+ Everything else β€” `DESIGN.md` (decisions and why), `CHALLENGES.md` (every bug and contaminant),
103
+ the ablation report, the fine-tuning log with corrections, and every evaluated generation β€” is in the
104
+ GitHub repository. MIT licensed.
SHA256SUMS.txt ADDED
@@ -0,0 +1 @@
 
 
1
+ d2c030b0fbd8c9a6a227b96e704a8c538c3754f468b2ebd2d9970f880b8c7b0d model.safetensors
arch_facts.json ADDED
@@ -0,0 +1,29 @@
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "params_total": 434680832,
3
+ "n_tensors": 157,
4
+ "dtype": "bf16",
5
+ "vocab_size": null,
6
+ "d_model": null,
7
+ "n_blocks": 22,
8
+ "tensor_keys_sample": [
9
+ "blocks.0.attn.out_proj.weight",
10
+ "blocks.0.attn.qkv.weight",
11
+ "blocks.0.attn_norm.scale",
12
+ "blocks.0.ffn.down.weight",
13
+ "blocks.0.ffn.gate.weight",
14
+ "blocks.0.ffn.up.weight",
15
+ "blocks.0.ffn_norm.scale",
16
+ "blocks.1.attn.out_proj.weight"
17
+ ],
18
+ "source_checkpoint": "the pretrained checkpoint this export was taken from (path omitted in the released copy)",
19
+ "source_keys": [
20
+ "optimizer",
21
+ "scheduler",
22
+ "config",
23
+ "training_state",
24
+ "timestamp",
25
+ "model"
26
+ ],
27
+ "step": null,
28
+ "created_utc": "2026-09-20T18:14:35Z"
29
+ }
config.json ADDED
@@ -0,0 +1,10 @@
 
 
 
 
 
 
 
 
 
 
 
1
+ {
2
+ "vocab_size": 32000,
3
+ "d_model": 1024,
4
+ "num_layers": 22,
5
+ "num_heads": 16,
6
+ "d_ff": 4096,
7
+ "max_seq_len": 1024,
8
+ "rope_theta": 500000.0,
9
+ "dropout_rate": 0.0
10
+ }
model.safetensors ADDED
@@ -0,0 +1,3 @@
 
 
 
 
1
+ version https://git-lfs.github.com/spec/v1
2
+ oid sha256:d2c030b0fbd8c9a6a227b96e704a8c538c3754f468b2ebd2d9970f880b8c7b0d
3
+ size 869377400
tokenizer.json ADDED
The diff for this file is too large to render. See raw diff