slm125m-base

A 126M-parameter Llama-architecture base model pretrained from scratch on a legal/financial corpus. Base model — not instruction-tuned.

Architecture

Parameters 125,848,320
Layers / hidden / heads 12 / 768 / 12
Context length 1024
Vocab 16384 (byte-level BPE trained from scratch on this corpus)
Tied embeddings True

Training data

~2.19B tokens, deduplicated and decontaminated. Mix: case-law ~39% / SEC ~39% / fineweb-edu ~21%.

Sources: HFforLegal/case-law (US court opinions), PleIAs/SEC (SEC filings), HuggingFaceFW/fineweb-edu (sample-10BT, fluency filler).

Pipeline: 6-step deterministic cleaning (line filter, boilerplate strip, repetition, English gate, OCR-garble gate on case-law) then exact-hash dedup, MinHash/LSH near-dup removal, and 13-gram decontamination against CaseHOLD / LexGLUE case_hold, which are therefore held out.

Training

8.28B tokens seen (4 epochs over the corpus), 15,789 steps, global batch 524,288 tokens, AdamW, cosine 0.0006 → 6e-05, bf16 on 8×H100.

Results

Validation loss 2.0934 (perplexity 8.11) on a held-out 1% split of the same corpus. This is in-domain perplexity and is not comparable across models trained on different data or tokenizers.

Limitations

Small base model. It will produce fluent-looking but frequently incorrect legal and financial text, and it has no instruction tuning, no alignment, and no factual grounding. Not legal or financial advice; do not use for either.

Downloads last month
15
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support