theOG-50M
An experimental PyTorch checkpoint from a small language model research project (compute-elastic 50M-class pilot).
Files
model_state.ptโ raw state dict (torch.load(..., weights_only=True), fp32, ~208 MB).config.jsonโ architecture: 16000 vocab, hidden 512, 10 layers, 4 heads x head_dim 128, SwiGLU, RoPE + learned positional embeddings, tied embeddings, MoE with 2 experts (top_k=1).tokenizer.jsonโ BPE tokenizer (15853 tokens) matching the checkpoint. Reserved ids:<|think|>= 15851,<|/think|>= 15852.
Size
Effective model size is ~50M active parameters (50.7M per the
pilot_50m_deep config: MoE with 2 experts, only 1 expert active per token,
embeddings tied). The raw .pt state dict totals 60.2M values because it
counts the tied embedding twice and includes both dormant experts.
Load and run
import torch
from tokenizers import Tokenizer
sd = torch.load("model_state.pt", map_location="cpu", weights_only=True)
tok = Tokenizer.from_file("tokenizer.json")
No reference implementation is included in this repo, so the exact wiring must
be reconstructed from the tensor shapes in config.json + model_state.pt.
HumanEval (50-task subset, execution-graded)
Measured on the project's canonical HumanEval harness: sample k completions
(temp 0.8, top-p 0.95, max 220 new tokens), execution-select the first passing
one. Pass iff prompt + completion + test + check(entry_point) exits 0.
| protocol | pass@1 | pass@k |
|---|---|---|
| k=8 | 11/50 | 16/50 |
| k=16 | 11/50 | 19/50 |
| k=32 | 11/50 | 20/50 |
| CodeGen-350M greedy (reference) | 15/50 | 15/50 |
The full per-task completions and grader are in the project repo
(scripts/bench_humaneval.py, evaluation/results/humaneval_og50_k{8,16,32}.json).
Status
Experimental checkpoint from a research run. The numbers above are the measured ones with the protocol stated; single-sample pass@1 is significant (11/50), sampling with execution selection brings it to ~16-20/50.
Safety note
The state dict was disassembled before loading; it contains only standard
tensor-rebuild globals and is safe to load with weights_only=True.
- Downloads last month
- -