Download README.md from Compactbot/ldt-10m: direct link, hf CLI and curl.
- Browser
- Download file 3.59 kB
-
https://huggingface.co/Compactbot/ldt-10m/resolve/refs%2Fpr%2F7/README.md
- Command line
-
hf download hf://Compactbot/ldt-10m@refs/pr/7/README.md
-
curl -L -o README.md https://huggingface.co/Compactbot/ldt-10m/resolve/refs%2Fpr%2F7/README.md
license: apache-2.0
pipeline_tag: text-generation
language: en
tags:
- tiny
- tiny-lm
- tiny-model
- slm
- SLM
- small-language-model
- from-scratch
- llama
datasets:
- HuggingFaceFW/fineweb-edu
- allenai/dclm-baseline
metrics:
- perplexity
LDT-10M
A 10.28M-parameter LLaMA-style language model, trained from scratch on FineWeb-Edu + DCLM.
Requested by DedeProGames on the model-requests board (#12).
Architecture
| Parameter | Value |
|---|---|
| Params | 10,284,480 |
| Layers | 5 |
| d_model | 320 |
| Heads | 5 (MHA, GQA not used at this scale) |
| FFN dim | 896 (SwiGLU) |
| Vocab | 12,288 (gollem BPE) |
| Context | 512 |
| Embeddings | Tied (lm_head → tok.weight) |
| Norm | RMSNorm (eps 1e-5) |
| Attention | RoPE + causal SDPA |
| Dtype | float32 |
Standard LLaMA block: RMSNorm → MHA (RoPE) → residual → RMSNorm → SwiGLU FFN → residual.
Training
| Steps | 4,000 |
| Batch size | 64 |
| LR | 3e-4 → 3e-5 (cosine decay) |
| Data | FineWeb-Edu (30.1M tok) + DCLM baseline (21.4M tok) = ~50.5M tokens |
| Tokens/param | ~4.9 |
| Hardware | RTX 5090 (32 GB), GPU |
| Final val loss | 4.6020 (ppl 99.68) |
⚠️ Honest caveat: undertrained
DedeProGames requested 2.6B tokens. This checkpoint is at 50.5M tokens — a ~50× shortfall. The GPU was occupied by other work for most of the training window, and the CPU was over-subscribed.
At 4.9 tok/param, the model has learned grammar and surface fluency but not deep coherence. The val loss (4.602) is well below the 7.38 unigram floor, so it genuinely uses context — but the prose is semantically thin: wordy, repetitive, and it drifts off-topic mid-sentence.
This is a first checkpoint, not the final deliverable. Continued training toward the 2.6B budget is planned.
Eval (40 samples, 8 prompts × 5 seeds, temp 0.8, top-k 40)
| Metric | Value |
|---|---|
| mean loop_frac | 0.042 |
| max loop_frac | 1.0 (one sample) |
| Degenerate? | No |
| Below unigram floor? | Yes (4.602 < 7.38) |
Sample outputs
"The sun is assembled by a new study because he is an associate of the study of the disease in the early years. I've been interested in having a very different study…"
"Once upon a time, she is an attack. But he is not a good idea. But it is something that does not have a moment of his own life…"
"The cat sat on the ground. The sunp is a piece of light and is not a good deal of tear. The new story of the MD's Ford…"
"def hello(): I have a lot of the best. I'm not sure what happened to me. I'll be able to do anything, but I think it's been just a lot to say…"
Grammar is intact. No token loops, no broken tokens, no speaker tags. Semantically limited — expected at 4.9 tok/param.
Usage
The model uses a custom LDT architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares tok.weight).
To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (train_ldt10m_fixed.py) contains the full architecture definition.
What this is NOT
- Not a 2.6B-token model (that's the target; this is the 50.5M checkpoint)
- Not a general-purpose assistant (it's a raw LM, no instruction tuning)
- Not a replacement for anything larger — it's a research checkpoint in a from-scratch training run