Text Generation
MLX
Safetensors
English
pretraining
from-scratch
small-language-model
post-training
silicon
Instructions to use OpenSML/OpenSML-150M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use OpenSML/OpenSML-150M with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # if on a CUDA device, also pip install mlx[cuda] # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("OpenSML/OpenSML-150M") prompt = "Once upon a time in" text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- MLX LM
How to use OpenSML/OpenSML-150M with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Generate some text mlx_lm.generate --model "OpenSML/OpenSML-150M" --prompt "Once upon a time"
- Atomic Chat
File size: 11,542 Bytes
8662ab2 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 240 241 242 243 244 245 246 247 248 249 250 251 252 253 254 255 256 257 258 259 260 261 262 263 264 265 266 267 268 269 270 271 272 273 274 275 276 277 278 279 280 281 282 283 | ---
license: apache-2.0
datasets:
- HuggingFaceTB/smollm-corpus
- HuggingFaceTB/dclm-edu
- HuggingFaceFW/finewiki
- HuggingFaceTB/smol-smoltalk
- HuggingFaceH4/ultrachat_200k
- rajpurkar/squad_v2
- allenai/ai2_arc
- databricks/databricks-dolly-15k
- allenai/tulu-3-sft-personas-instruction-following
- HuggingFaceTB/smoltalk
- allenai/sciq
language:
- en
pipeline_tag: text-generation
tags:
- text-generation
- mlx
- pretraining
- from-scratch
- small-language-model
- post-training
- silicon
---
# OpenSML-150M
**Author:** William Zebrowski · **Checkpoint:** OpenSML-150M Instruct · **Status:** research preview
[Technical report](https://github.com/williamzebrowskI/opensml-150m/blob/main/sml-mlx-v1/docs/TECHNICAL_REPORT.md) · [Results and provenance](TRIAL28_RESULTS.json) · [Tokenizer](tokenizer/tokenizer.json)
An English-first language model trained from scratch with Apple's MLX framework.
Pretraining processed **7.800B tokens** across Stage A and Stage B on a four-Mac
cluster, later expanded to five Macs using Thunderbolt RDMA. Three supervised
fine-tuning stages produced the selected OpenSML-150M checkpoint.
## Availability and Inference
The selected **OpenSML-150M** weights, frozen tokenizer, and standalone native MLX inference code are included. The **OpenSML-150M Instruct** checkpoint contains the selected fine-tuned FP32 weights; optimizer state is excluded. Its training-step identity is retained in the checkpoint provenance.
Use Python 3.11+ on Apple Silicon. Install the Hugging Face CLI first and authenticate if this repository is private:
```bash
hf download wzebrowski/OpenSML-150M --local-dir ./OpenSML-150M
cd OpenSML-150M
python -m pip install -r requirements.txt
python inference.py --prompt "Say hello in one sentence." --max-new-tokens 32
```
The CLI verifies the weight hash and tokenizer manifest, then uses greedy FP32 reference MLX inference. Prompts use `User: {prompt}\nAssistant:`; `--raw-completion` skips this wrapper. The 2,048-token context includes the requested generation budget; overflow is rejected. Output includes text, generated token IDs, and the stop reason. Native loading and a short generation were verified locally; this is **not a Transformers AutoModel or mlx-lm loader package**. Other formats and cross-framework parity remain unverified.
See [checkpoint provenance](checkpoint_provenance.json), [loading verification](INFERENCE_VERIFICATION.json), and [bundle hashes](SHA256SUMS).
## Model Specification
| Field | Value |
| --- | --- |
| Parameters (model-card label) | 150,439,188 (150.44M) |
| Layers / hidden width / FFN width | 20 / 768 / 2,048 |
| Attention | GQA: 12 query heads / 4 KV heads; head dimension 64 |
| Position encoding | RoPE, base 10,000 |
| Normalization / MLP | RMSNorm, Q/K normalization, SwiGLU |
| Vocabulary / context | 32,000 / 2,048 |
| Embeddings | Tied input and output |
| Linear biases / dropout | Neither used |
| Pretraining precision | BF16 compute; FP32 master weights and optimizer state |
| Benchmark precision | FP32 scoring, explicit vanilla MLX attention |
The independently fitted tokenizer is a 32,000-token byte-level BPE. Its recorded
held-out audit had zero round-trip failures. Full configuration, parameter
accounting, and tokenizer checks are in the [technical report](https://github.com/williamzebrowskI/opensml-150m/blob/main/sml-mlx-v1/docs/TECHNICAL_REPORT.md).
## Training and Model Lineage
The selected pretrained base is **Stage B step 73,243**, after **7,800,086,528
lifetime pretraining tokens**. Stage B continued Stage A with the adjusted mixture below.
Percentages are configured token shares.
| Source | Initial token share | Stage B token share | Selection |
| --- | ---: | ---: | --- |
| [FineWeb-Edu](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) | 55% | 55% | `fineweb-edu-dedup` |
| [DCLM-Edu](https://huggingface.co/datasets/HuggingFaceTB/dclm-edu) | 25% | 20% | `edu_int_score >= 3` |
| [FineWiki](https://huggingface.co/datasets/HuggingFaceFW/finewiki) | 10% | 15% | English |
| [Cosmopedia v2](https://huggingface.co/datasets/HuggingFaceTB/smollm-corpus) | 10% | 10% | Textbook/tutorial/blog/educational format filter |
FineWeb-Edu and Cosmopedia v2 come from two subsets of the same SmolLM Corpus
repository. Pretraining used no dedicated code or math dataset.
```text
V1 pretrained base → Unified384 → Repair512 → OpenSML-150M (SFT step 768)
+384 +128 +256 updates
```
Unified384 used Smol-SmolTalk, UltraChat, SQuAD v2, ARC training questions, and
locally authored follow-ups. Repair512 added Dolly and Tulu Persona instruction
examples. The final stage used SmolTalk constraints, SQuAD v2, SciQ, and Repair512
replay. All model parameters were updated; the selected checkpoint is a direct
continuation, without LoRA or parameter interpolation. Dataset exposure, source
pins, optimizer schedules, and checkpoint hashes are in the technical report.
### Pretraining Loss Curves

Recorded held-out loss without smoothing. The star marks the selected base;
later points are outside its training exposure. The late-training panel uses a
magnified scale. This chart uses the same original validation set throughout.
<details>
<summary>Stage B detail and per-source curves</summary>

This panel uses a different, expanded validation set; its absolute losses should
not be merged with the overview. The selected base has original-set loss
2.81855024 and expanded-set loss 2.76700753.

Per-source panels use separate vertical scales. Full validation accounting and
historical measurements are in the technical report.
</details>
## Evaluation
### Zero-shot Likelihood Benchmarks
Existing full-split evaluations cover **15,428 multiple-choice examples**.
The pretrained base and selected SFT model are reported separately.
| Benchmark | Split | Examples | Pretrained base acc / acc_norm | OpenSML-150M acc / acc_norm |
| --- | --- | ---: | ---: | ---: |
| ARC-Easy | Test | 2,376 | 54.67% / 48.53% | 56.65% / 55.43% |
| ARC-Challenge | Test | 1,172 | 23.38% / 26.88% | 26.02% / 29.52% |
| PIQA | Validation | 1,838 | 65.23% / 64.36% | 64.53% / 64.09% |
| HellaSwag | Validation | 10,042 | 30.21% / 34.44% | 30.55% / 33.94% |
SFT improves both ARC metrics. PIQA declines slightly; HellaSwag raw accuracy
rises slightly while normalized accuracy declines.
<details>
<summary>Likelihood prompts and scoring</summary>
The native evaluator uses zero-shot raw completion prompts, FP32 scoring,
reference MLX attention, batch size 1, and no chat template or added BOS/EOS.
`acc` ranks summed candidate log-likelihood; `acc_norm` divides by candidate
Unicode character count, excluding the leading delimiter. No generation or LLM
judge is involved. The evaluator follows pinned harness conventions but is not
an installed lm-evaluation-harness run. Dataset pins and integrity receipts are
in the technical report and results record.
</details>
### Comparison with Small Language Models
External scores below are **previously published measurements**. Each cell shows
`acc / acc_norm`; **—** means not reported in the selected source. Bold marks the
highest available reported value.
| Model | ARC-Easy acc / acc_norm | ARC-Challenge acc / acc_norm | PIQA acc / acc_norm | HellaSwag acc / acc_norm |
| --- | ---: | ---: | ---: | ---: |
| **OpenSML-150M (SFT)** | **56.65% / 55.43%** | **26.02% / 29.52%** | **64.53% / 64.09%** | **30.55% / 33.94%** |
| [GPT-2 (124M, base)](https://huggingface.co/openai-community/gpt2) | — / 39.48% | — / — | — / 62.51% | 28.92% / 31.14% |
| [OPT-125M (base)](https://huggingface.co/facebook/opt-125m) | 43.52% / 39.98% | 18.94% / 22.78% | 63.00% / 62.02% | — / — |
| [Pythia-160M (base)](https://huggingface.co/EleutherAI/pythia-160m) | 43.52% / 39.65% | 18.77% / 23.29% | 62.73% / 61.64% | — / — |
OpenSML's largest reported advantage is normalized ARC-Easy: +15.45 percentage
points over OPT-125M and +15.78 over Pythia-160M. These are cross-source
comparisons, not a controlled ranking: prompts, normalization, numerical settings,
and dataset revisions are not verified identical. OpenSML is SFT-trained; the
references are base models. Source records are in the technical report.
### Instruction Following: IFEval
| Checkpoint | Prompt strict | Instruction strict | Prompt loose | Instruction loose | 1,280-token cap hits |
| --- | ---: | ---: | ---: | ---: | ---: |
| OpenSML-150M | 15.16% (82/541) | 25.30% (211/834) | 15.71% (85/541) | 25.78% (215/834) | 33/541 |
These results apply to OpenSML-150M, not the pretrained base. All 541 prompts and
834 instructions were scored using programmatic strict/loose verifiers, without
an LLM judge.
<details>
<summary>IFEval formatting, generation, and scoring</summary>
Generation is zero-shot and greedy, with plain `User: {prompt}\nAssistant:`
formatting, no added system message, and a 1,280-new-token cap. The context is
2,048 tokens; document-end EOS is ID 1. No additional stop strings or repetition
penalty are used. All responses are scored as produced: 508 stop at EOS and 33
reach the generation cap; none hit the context limit. The dataset's only split
is named `train`, but these prompts are evaluation inputs, not supervised records.
</details>
### Complete-answer Diagnostics
Saved development checks report a follow-up joint proxy of 22/32 and named-constraint
passes on 5/32 public cases. Repetition was detected on 4/63 legacy turns and
18/73 public-development turns. These are mechanical development proxies;
a systematic independent review of complete-answer correctness is still pending.
Likelihood and IFEval do not establish reliable factual answers. No MT-Bench or
coding-success score is reported for the selected model.
## Intended Use and Limitations
Intended for small-language-model research and controlled experimentation.
Responses may be incorrect, repetitive, biased, or unsafe. Production use,
safety behavior, tool use, multilingual ability, and extended dialogue remain
unvalidated. Benchmark reuse during selection limits evaluation independence;
a comprehensive contamination audit has not been established.
## Checkpoint Verification
Saved hashes and integrity receipts identify the selected weights and completed
evaluations. They do not establish cross-framework inference parity. An unchanged native weight export and standalone inference CLI are included.
The export has a local loading/generation check; a full benchmark rerun of this download bundle has not been performed. See the
[results and provenance record](TRIAL28_RESULTS.json) and technical report.
## Licensing and Data Provenance
William Zebrowski releases the OpenSML-150M model weights, tokenizer and author-controlled project contributions under [Apache-2.0](LICENSE). Anyone may use, modify, fine-tune and redistribute these contributions for personal, research or commercial purposes, subject to the license terms.
This grant covers rights held by the author and does not relicense third-party training materials. Upstream materials retain their own terms, including SQuAD v2 (CC-BY-SA-4.0) and SciQ (CC-BY-NC-3.0). The model license does not grant blanket commercial rights to those datasets.
|