--- license: apache-2.0 library_name: transformers pipeline_tag: text-generation tags: - llama - small-language-model - pretrained-from-scratch - legal --- # slm-125m-base A **125.8M-parameter** Llama-shaped language model pretrained **from scratch** on a legal-first corpus (US case law + SEC filings + a slice of FineWeb-Edu). Built with a custom 16,384-token BPE tokenizer. This is a **base / foundation** model -- it does next-token continuation, not instruction following or chat. Expect rough, domain-flavored completions; it was trained on a small budget (step 19,334, ~10.14B tokens seen). ## Architecture | | | |---|---| | Params | ~125.8M (tied embeddings) | | Layers | 12 | | Hidden size | 768 | | Heads / KV heads | 12 / 12 (MHA) | | Context length | 1,024 | | Vocab | 16,384 (custom BPE) | | Activation | SwiGLU (silu) | | Position | RoPE (theta 10000) | ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer import torch tok = AutoTokenizer.from_pretrained("analyticspro/slm-125m-base") model = AutoModelForCausalLM.from_pretrained("analyticspro/slm-125m-base") model.eval() ids = tok("The court held that", return_tensors="pt") out = model.generate(**ids, max_new_tokens=120, do_sample=True, temperature=0.8, top_p=0.95, pad_token_id=tok.eos_token_id) print(tok.decode(out[0], skip_special_tokens=True)) ``` ## Limitations Small model, small pretraining budget, and a legal-heavy corpus: outputs can be factually wrong, repetitive, or biased toward legal/regulatory phrasing. Not suitable for production or any high-stakes use. For research and demos only.