slm-125m-base

A 125M-parameter, Llama-style, decoder-only language model, pretrained completely from scratch (fresh weights, fresh 16K byte-level BPE tokenizer) on a domain-specific corpus of US case law, SEC filings, and general web text.

This is a base completer, not a chat model โ€” give it a passage prefix and it continues it. It has not been instruction-tuned, aligned, or fine-tuned for any downstream task.

Architecture

Parameters ~125M
Layers 12
Hidden size 768
Attention heads 12 (head dim 64), plain MHA
MLP SwiGLU, inner 3072
Normalization RMSNorm, pre-norm
Position embeddings RoPE
Context length 1024
Vocab 16,384 (fresh byte-level BPE, no OOV token ever)
Embeddings tied

Training data

~3.53B packed tokens across three sources, deduplicated (MinHash-LSH on the legal corpus) and decontaminated against the LexGLUE/CaseHOLD held-out splits:

Source Tokens Share
HFforLegal/case-law (US case law, us split) ~1.1B 31%
PleIAs/SEC (SEC filings) ~2.0B 57%
HuggingFaceFW/fineweb-edu (general web text) ~1.0B 28%

(Shares don't sum to 100% due to rounding; case-law's split was capped at its full-dataset ceiling, not a target percentage โ€” see the build's docs/01-data.md.)

Training

AdamW (ฮฒ 0.9/0.95, weight decay 0.1), linear warmup + cosine LR, bf16 (autocast, fp32 master weights), single-node DDP across 8x H100. Trained to step 7,500 (~4.1B tokens processed, more than one full epoch over the training set) before the run was stopped.

Evaluation

Held-out validation perplexity: 10.14 (full validation set, 34,513 windows). This is the honest headline metric at this scale โ€” not MMLU or similar benchmarks, which are near-random for a 125M model.

Perplexity varies a lot by domain, reflecting the training mix:

Domain Perplexity
SEC filings ~6.0
Case law ~15.3
General web text (fineweb-edu) ~27.4

The model is noticeably better at the two legal/financial domains it mostly trained on than at general web text.

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("mickey5k/slm-125m-base")
model = AutoModelForCausalLM.from_pretrained("mickey5k/slm-125m-base")

prompt = "The court held that"
input_ids = tokenizer(prompt, return_tensors="pt").input_ids
output_ids = model.generate(
    input_ids,
    max_new_tokens=100,
    min_new_tokens=20,  # a base model can emit EOS almost immediately otherwise
    do_sample=True,
    temperature=0.8,
    top_p=0.9,
    top_k=50,
)
print(tokenizer.decode(output_ids[0], skip_special_tokens=True))

Limitations

  • Base completer, not instruction-tuned. It continues text; it does not follow instructions or answer questions in a chat format.
  • Domain-skewed. Strongest on legal/financial register, weaker on general text (see per-domain perplexity above).
  • Small. 125M parameters is far below the scale where benchmark-style reasoning emerges โ€” evaluate it on perplexity/completion quality, not accuracy benchmarks.
  • Training was cost-capped, not run to convergence โ€” see the build's docs/05-pretrain.md for the full training log and the reasoning behind stopping at step 7,500.
Downloads last month
40
Safetensors
Model size
0.1B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support