theOG-50M

An experimental PyTorch checkpoint from a small language model research project (compute-elastic 50M-class pilot).

Files

  • model_state.pt โ€” raw state dict (torch.load(..., weights_only=True), fp32, ~208 MB).
  • config.json โ€” architecture: 16000 vocab, hidden 512, 10 layers, 4 heads x head_dim 128, SwiGLU, RoPE + learned positional embeddings, tied embeddings, MoE with 2 experts (top_k=1).
  • tokenizer.json โ€” BPE tokenizer (15853 tokens) matching the checkpoint. Reserved ids: <|think|> = 15851, <|/think|> = 15852.

Size

Effective model size is ~50M active parameters (50.7M per the pilot_50m_deep config: MoE with 2 experts, only 1 expert active per token, embeddings tied). The raw .pt state dict totals 60.2M values because it counts the tied embedding twice and includes both dormant experts.

Load and run

import torch
from tokenizers import Tokenizer

sd = torch.load("model_state.pt", map_location="cpu", weights_only=True)
tok = Tokenizer.from_file("tokenizer.json")

No reference implementation is included in this repo, so the exact wiring must be reconstructed from the tensor shapes in config.json + model_state.pt.

HumanEval (50-task subset, execution-graded)

Measured on the project's canonical HumanEval harness: sample k completions (temp 0.8, top-p 0.95, max 220 new tokens), execution-select the first passing one. Pass iff prompt + completion + test + check(entry_point) exits 0.

protocol pass@1 pass@k
k=8 11/50 16/50
k=16 11/50 19/50
k=32 11/50 20/50
CodeGen-350M greedy (reference) 15/50 15/50

The full per-task completions and grader are in the project repo (scripts/bench_humaneval.py, evaluation/results/humaneval_og50_k{8,16,32}.json).

Status

Experimental checkpoint from a research run. The numbers above are the measured ones with the protocol stated; single-sample pass@1 is significant (11/50), sampling with execution selection brings it to ~16-20/50.

Safety note

The state dict was disassembled before loading; it contains only standard tensor-rebuild globals and is safe to load with weights_only=True.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support