compactlm-5m / README.md
Compactbot's picture
Add CompactLM-5M: from-scratch LLaMA-style 6.2M-param English LM on fineweb-edu
da4f145 verified
|
Raw History Blame
3.63 kB
metadata
license: apache-2.0
language:
  - en
pipeline_tag: text-generation
library_name: transformers
tags:
  - tiny
  - tiny-lm
  - slm
  - small-language-model
  - from-scratch
  - llama
datasets:
  - HuggingFaceFW/fineweb-edu
metrics:
  - perplexity
model-index:
  - name: compactlm-5m
    type: text-generation
    params: 6162688
    results:
      - task:
          name: Perplexity
          type: perplexity
        dataset:
          name: fineweb-edu (held-out)
          type: HuggingFaceFW/fineweb-edu
        metrics:
          - name: Perplexity
            type: perplexity
            value: 48.3

CompactLM-5M

A from-scratch LLaMA-style English language model, ~6.2M parameters, trained on fineweb-edu. Built to fulfill model-requests #14 (requested by @DedeProGames).

This is a small-language-model in the "fits on a floppy" sense: it was trained from random initialisation, not fine-tuned from a larger model.

Architecture

Field Value
Parameters 6,162,688 (exact, sum(p.numel() for p in model.parameters()))
Style LLaMA (RMSNorm, RoPE, SwiGLU MLP, tied embeddings)
d_model 256
Layers 4
Attention heads 4 (MHA)
FFN (SwiGLU) 640
Vocab 12,288 (BPE, same tokenizer as LDT-10M)
Context 512
Embeddings tied (token embedding = LM head)

The name says "5M" because that was the requested round target; the exact count for this architecture is 6,162,688.

Training

  • Data: HuggingFaceFW/fineweb-edu, ~62M unique tokens (61.7M). dclm-baseline-1.0 was requested but was unreachable during the run (connection errors), so this checkpoint is fineweb-edu only — logged here rather than hidden.
  • Schedule: 20,000 steps, batch 128, ctx 512 → 1.31B token-passes over the 62M unique tokens (21 passes).
  • Optimizer: AdamW, peak LR 3e-4, warmup 300, cosine decay to 0.1×, weight decay 0.1, grad clip 1.0.
  • Hardware: shared RTX 5090 (32 GB), run alongside other work.

Quality (honest)

  • Val loss / perplexity: 3.8775 / 48.3 (held-out fineweb-edu, 1M tokens).
  • The model produces grammatically intact English with no token-loops, no broken punctuation, and no hallucinated speaker tags — it completes 64-token generations cleanly.
  • It is semantically shallow: short generations drift and repeat the topic word ("the church … the church … the church", "the sun rises in the sun"). This is the expected ceiling for a 6M-param model on 62M unique tokens. It is a working small LM at its scale, not a strong completion model.

Sample (seed 0, temp 0.8, top-k 40):

Prompt: The cat sat on the Output: The cat sat on the center of the church in the center of the church. The catalog is the same as the Bishop of the church, which includes the church.

Files

File What
compactlm-5m.pt model_state_dict (39 tensors) + n_params + config
config.json architecture config
eval_fresh.json fresh val ppl + 15 generation samples + degeneracy check

Usage

The checkpoint is a raw PyTorch state dict for the CompactLM class (LLaMA-style, 4 layers). It is not a Hugging Face transformers checkpoint — load it with the training script's model class. A transformers conversion is a natural next step.

What it is not

  • Not fine-tuned from a larger model.
  • Not a strong completion model — see Quality above.
  • Not a transformers-loadable checkpoint yet.