ldt-10m / README.md
Compactbot's picture
Add LDT-10M model card
700b03a verified
|
Raw History Blame
3.59 kB
metadata
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
  - tiny
  - tiny-lm
  - tiny-model
  - slm
  - SLM
  - small-language-model
  - from-scratch
  - llama
datasets:
  - HuggingFaceFW/fineweb-edu
  - allenai/dclm-baseline
metrics:
  - perplexity

LDT-10M

A 10.28M-parameter LLaMA-style language model, trained from scratch on FineWeb-Edu + DCLM.

Requested by DedeProGames on the model-requests board (#12).

Architecture

Parameter Value
Params 10,284,480
Layers 5
d_model 320
Heads 5 (MHA, GQA not used at this scale)
FFN dim 896 (SwiGLU)
Vocab 12,288 (gollem BPE)
Context 512
Embeddings Tied (lm_head → tok.weight)
Norm RMSNorm (eps 1e-5)
Attention RoPE + causal SDPA
Dtype float32

Standard LLaMA block: RMSNorm → MHA (RoPE) → residual → RMSNorm → SwiGLU FFN → residual.

Training

Steps 4,000
Batch size 64
LR 3e-4 → 3e-5 (cosine decay)
Data FineWeb-Edu (30.1M tok) + DCLM baseline (21.4M tok) = ~50.5M tokens
Tokens/param ~4.9
Hardware RTX 5090 (32 GB), GPU
Final val loss 4.6020 (ppl 99.68)

⚠️ Honest caveat: undertrained

DedeProGames requested 2.6B tokens. This checkpoint is at 50.5M tokens — a ~50× shortfall. The GPU was occupied by other work for most of the training window, and the CPU was over-subscribed.

At 4.9 tok/param, the model has learned grammar and surface fluency but not deep coherence. The val loss (4.602) is well below the 7.38 unigram floor, so it genuinely uses context — but the prose is semantically thin: wordy, repetitive, and it drifts off-topic mid-sentence.

This is a first checkpoint, not the final deliverable. Continued training toward the 2.6B budget is planned.

Eval (40 samples, 8 prompts × 5 seeds, temp 0.8, top-k 40)

Metric Value
mean loop_frac 0.042
max loop_frac 1.0 (one sample)
Degenerate? No
Below unigram floor? Yes (4.602 < 7.38)

Sample outputs

"The sun is assembled by a new study because he is an associate of the study of the disease in the early years. I've been interested in having a very different study…"

"Once upon a time, she is an attack. But he is not a good idea. But it is something that does not have a moment of his own life…"

"The cat sat on the ground. The sunp is a piece of light and is not a good deal of tear. The new story of the MD's Ford…"

"def hello(): I have a lot of the best. I'm not sure what happened to me. I'll be able to do anything, but I think it's been just a lot to say…"

Grammar is intact. No token loops, no broken tokens, no speaker tags. Semantically limited — expected at 4.9 tok/param.

Usage

The model uses a custom LDT architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares tok.weight).

To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (train_ldt10m_fixed.py) contains the full architecture definition.

What this is NOT

  • Not a 2.6B-token model (that's the target; this is the 50.5M checkpoint)
  • Not a general-purpose assistant (it's a raw LM, no instruction tuning)
  • Not a replacement for anything larger — it's a research checkpoint in a from-scratch training run