File size: 11,542 Bytes
8662ab2
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
235
236
237
238
239
240
241
242
243
244
245
246
247
248
249
250
251
252
253
254
255
256
257
258
259
260
261
262
263
264
265
266
267
268
269
270
271
272
273
274
275
276
277
278
279
280
281
282
283
---
license: apache-2.0
datasets:
- HuggingFaceTB/smollm-corpus
- HuggingFaceTB/dclm-edu
- HuggingFaceFW/finewiki
- HuggingFaceTB/smol-smoltalk
- HuggingFaceH4/ultrachat_200k
- rajpurkar/squad_v2
- allenai/ai2_arc
- databricks/databricks-dolly-15k
- allenai/tulu-3-sft-personas-instruction-following
- HuggingFaceTB/smoltalk
- allenai/sciq
language:
- en
pipeline_tag: text-generation
tags:
- text-generation
- mlx
- pretraining
- from-scratch
- small-language-model
- post-training
- silicon

---


# OpenSML-150M


**Author:** William Zebrowski · **Checkpoint:** OpenSML-150M Instruct · **Status:** research preview


[Technical report](https://github.com/williamzebrowskI/opensml-150m/blob/main/sml-mlx-v1/docs/TECHNICAL_REPORT.md) · [Results and provenance](TRIAL28_RESULTS.json) · [Tokenizer](tokenizer/tokenizer.json)


An English-first language model trained from scratch with Apple's MLX framework.
Pretraining processed **7.800B tokens** across Stage A and Stage B on a four-Mac
cluster, later expanded to five Macs using Thunderbolt RDMA. Three supervised
fine-tuning stages produced the selected OpenSML-150M checkpoint.


## Availability and Inference


The selected **OpenSML-150M** weights, frozen tokenizer, and standalone native MLX inference code are included. The **OpenSML-150M Instruct** checkpoint contains the selected fine-tuned FP32 weights; optimizer state is excluded. Its training-step identity is retained in the checkpoint provenance.


Use Python 3.11+ on Apple Silicon. Install the Hugging Face CLI first and authenticate if this repository is private:


```bash
hf download wzebrowski/OpenSML-150M --local-dir ./OpenSML-150M
cd OpenSML-150M
python -m pip install -r requirements.txt
python inference.py --prompt "Say hello in one sentence." --max-new-tokens 32
```


The CLI verifies the weight hash and tokenizer manifest, then uses greedy FP32 reference MLX inference. Prompts use `User: {prompt}\nAssistant:`; `--raw-completion` skips this wrapper. The 2,048-token context includes the requested generation budget; overflow is rejected. Output includes text, generated token IDs, and the stop reason. Native loading and a short generation were verified locally; this is **not a Transformers AutoModel or mlx-lm loader package**. Other formats and cross-framework parity remain unverified.


See [checkpoint provenance](checkpoint_provenance.json), [loading verification](INFERENCE_VERIFICATION.json), and [bundle hashes](SHA256SUMS).


## Model Specification


| Field | Value |
| --- | --- |
| Parameters (model-card label) | 150,439,188 (150.44M) |
| Layers / hidden width / FFN width | 20 / 768 / 2,048 |
| Attention | GQA: 12 query heads / 4 KV heads; head dimension 64 |
| Position encoding | RoPE, base 10,000 |
| Normalization / MLP | RMSNorm, Q/K normalization, SwiGLU |
| Vocabulary / context | 32,000 / 2,048 |
| Embeddings | Tied input and output |
| Linear biases / dropout | Neither used |
| Pretraining precision | BF16 compute; FP32 master weights and optimizer state |
| Benchmark precision | FP32 scoring, explicit vanilla MLX attention |


The independently fitted tokenizer is a 32,000-token byte-level BPE. Its recorded
held-out audit had zero round-trip failures. Full configuration, parameter
accounting, and tokenizer checks are in the [technical report](https://github.com/williamzebrowskI/opensml-150m/blob/main/sml-mlx-v1/docs/TECHNICAL_REPORT.md).


## Training and Model Lineage


The selected pretrained base is **Stage B step 73,243**, after **7,800,086,528
lifetime pretraining tokens**. Stage B continued Stage A with the adjusted mixture below.
Percentages are configured token shares.


| Source | Initial token share | Stage B token share | Selection |
| --- | ---: | ---: | --- |
| [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) | 55% | 55% | `fineweb-edu-dedup` |
| [DCLM-Edu](https://huggingface.co/datasets/HuggingFaceTB/dclm-edu) | 25% | 20% | `edu_int_score >= 3` |
| [FineWiki](https://huggingface.co/datasets/HuggingFaceFW/finewiki) | 10% | 15% | English |
| [Cosmopedia v2](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) | 10% | 10% | Textbook/tutorial/blog/educational format filter |


FineWeb-Edu and Cosmopedia v2 come from two subsets of the same SmolLM Corpus
repository. Pretraining used no dedicated code or math dataset.


```text
V1 pretrained base → Unified384 → Repair512 → OpenSML-150M (SFT step 768)
                      +384         +128       +256 updates
```


Unified384 used Smol-SmolTalk, UltraChat, SQuAD v2, ARC training questions, and
locally authored follow-ups. Repair512 added Dolly and Tulu Persona instruction
examples. The final stage used SmolTalk constraints, SQuAD v2, SciQ, and Repair512
replay. All model parameters were updated; the selected checkpoint is a direct
continuation, without LoRA or parameter interpolation. Dataset exposure, source
pins, optimizer schedules, and checkpoint hashes are in the technical report.


### Pretraining Loss Curves


![Recorded pretraining validation loss for Stage A and Stage B](pretraining-overview.svg)


Recorded held-out loss without smoothing. The star marks the selected base;
later points are outside its training exposure. The late-training panel uses a
magnified scale. This chart uses the same original validation set throughout.


<details>
<summary>Stage B detail and per-source curves</summary>


![Stage B validation loss on the expanded held-out set](pretraining-phase-b.svg)


This panel uses a different, expanded validation set; its absolute losses should
not be merged with the overview. The selected base has original-set loss
2.81855024 and expanded-set loss 2.76700753.


![Original-set validation loss by data source](pretraining-sources.svg)


Per-source panels use separate vertical scales. Full validation accounting and
historical measurements are in the technical report.


</details>


## Evaluation


### Zero-shot Likelihood Benchmarks


Existing full-split evaluations cover **15,428 multiple-choice examples**.
The pretrained base and selected SFT model are reported separately.


| Benchmark | Split | Examples | Pretrained base acc / acc_norm | OpenSML-150M acc / acc_norm |
| --- | --- | ---: | ---: | ---: |
| ARC-Easy | Test | 2,376 | 54.67% / 48.53% | 56.65% / 55.43% |
| ARC-Challenge | Test | 1,172 | 23.38% / 26.88% | 26.02% / 29.52% |
| PIQA | Validation | 1,838 | 65.23% / 64.36% | 64.53% / 64.09% |
| HellaSwag | Validation | 10,042 | 30.21% / 34.44% | 30.55% / 33.94% |


SFT improves both ARC metrics. PIQA declines slightly; HellaSwag raw accuracy
rises slightly while normalized accuracy declines.


<details>
<summary>Likelihood prompts and scoring</summary>


The native evaluator uses zero-shot raw completion prompts, FP32 scoring,
reference MLX attention, batch size 1, and no chat template or added BOS/EOS.
`acc` ranks summed candidate log-likelihood; `acc_norm` divides by candidate
Unicode character count, excluding the leading delimiter. No generation or LLM
judge is involved. The evaluator follows pinned harness conventions but is not
an installed lm-evaluation-harness run. Dataset pins and integrity receipts are
in the technical report and results record.


</details>


### Comparison with Small Language Models


External scores below are **previously published measurements**. Each cell shows
`acc / acc_norm`; **—** means not reported in the selected source. Bold marks the
highest available reported value.


| Model | ARC-Easy acc / acc_norm | ARC-Challenge acc / acc_norm | PIQA acc / acc_norm | HellaSwag acc / acc_norm |
| --- | ---: | ---: | ---: | ---: |
| **OpenSML-150M (SFT)** | **56.65% / 55.43%** | **26.02% / 29.52%** | **64.53% / 64.09%** | **30.55% / 33.94%** |
| [GPT-2 (124M, base)](https://huggingface.co/openai-community/gpt2) | — / 39.48% | — / — | — / 62.51% | 28.92% / 31.14% |
| [OPT-125M (base)](https://huggingface.co/facebook/opt-125m) | 43.52% / 39.98% | 18.94% / 22.78% | 63.00% / 62.02% | — / — |
| [Pythia-160M (base)](https://huggingface.co/EleutherAI/pythia-160m) | 43.52% / 39.65% | 18.77% / 23.29% | 62.73% / 61.64% | — / — |


OpenSML's largest reported advantage is normalized ARC-Easy: +15.45 percentage
points over OPT-125M and +15.78 over Pythia-160M. These are cross-source
comparisons, not a controlled ranking: prompts, normalization, numerical settings,
and dataset revisions are not verified identical. OpenSML is SFT-trained; the
references are base models. Source records are in the technical report.


### Instruction Following: IFEval


| Checkpoint | Prompt strict | Instruction strict | Prompt loose | Instruction loose | 1,280-token cap hits |
| --- | ---: | ---: | ---: | ---: | ---: |
| OpenSML-150M | 15.16% (82/541) | 25.30% (211/834) | 15.71% (85/541) | 25.78% (215/834) | 33/541 |


These results apply to OpenSML-150M, not the pretrained base. All 541 prompts and
834 instructions were scored using programmatic strict/loose verifiers, without
an LLM judge.


<details>
<summary>IFEval formatting, generation, and scoring</summary>


Generation is zero-shot and greedy, with plain `User: {prompt}\nAssistant:`
formatting, no added system message, and a 1,280-new-token cap. The context is
2,048 tokens; document-end EOS is ID 1. No additional stop strings or repetition
penalty are used. All responses are scored as produced: 508 stop at EOS and 33
reach the generation cap; none hit the context limit. The dataset's only split
is named `train`, but these prompts are evaluation inputs, not supervised records.


</details>


### Complete-answer Diagnostics


Saved development checks report a follow-up joint proxy of 22/32 and named-constraint
passes on 5/32 public cases. Repetition was detected on 4/63 legacy turns and
18/73 public-development turns. These are mechanical development proxies;
a systematic independent review of complete-answer correctness is still pending.
Likelihood and IFEval do not establish reliable factual answers. No MT-Bench or
coding-success score is reported for the selected model.


## Intended Use and Limitations


Intended for small-language-model research and controlled experimentation.
Responses may be incorrect, repetitive, biased, or unsafe. Production use,
safety behavior, tool use, multilingual ability, and extended dialogue remain
unvalidated. Benchmark reuse during selection limits evaluation independence;
a comprehensive contamination audit has not been established.


## Checkpoint Verification


Saved hashes and integrity receipts identify the selected weights and completed
evaluations. They do not establish cross-framework inference parity. An unchanged native weight export and standalone inference CLI are included.
The export has a local loading/generation check; a full benchmark rerun of this download bundle has not been performed. See the
[results and provenance record](TRIAL28_RESULTS.json) and technical report.


## Licensing and Data Provenance


William Zebrowski releases the OpenSML-150M model weights, tokenizer and author-controlled project contributions under [Apache-2.0](LICENSE). Anyone may use, modify, fine-tune and redistribute these contributions for personal, research or commercial purposes, subject to the license terms.

This grant covers rights held by the author and does not relicense third-party training materials. Upstream materials retain their own terms, including SQuAD v2 (CC-BY-SA-4.0) and SciQ (CC-BY-NC-3.0). The model license does not grant blanket commercial rights to those datasets.