--- license: apache-2.0 pipeline_tag: text-generation language: en tags: - tiny - tiny-lm - tiny-model - slm - SLM - small-language-model - from-scratch - llama datasets: - HuggingFaceFW/fineweb-edu - allenai/dclm-baseline metrics: - perplexity --- # LDT-10M A 10.28M-parameter LLaMA-style language model, trained from scratch on FineWeb-Edu + DCLM. **Requested by [DedeProGames](https://huggingface.co/DedeProGames) on the [model-requests board](https://huggingface.co/spaces/Compactbot/model-requests) (#12).** ## Architecture | Parameter | Value | |---|---| | Params | **10,284,480** | | Layers | 5 | | d_model | 320 | | Heads | 5 (MHA, GQA not used at this scale) | | FFN dim | 896 (SwiGLU) | | Vocab | 12,288 (gollem BPE) | | Context | 512 | | Embeddings | Tied (lm_head → tok.weight) | | Norm | RMSNorm (eps 1e-5) | | Attention | RoPE + causal SDPA | | Dtype | float32 | Standard LLaMA block: RMSNorm → MHA (RoPE) → residual → RMSNorm → SwiGLU FFN → residual. ## Training | | | |---|---| | Steps | 4,000 | | Batch size | 64 | | LR | 3e-4 → 3e-5 (cosine decay) | | Data | FineWeb-Edu (30.1M tok) + DCLM baseline (21.4M tok) = **~50.5M tokens** | | Tokens/param | **~4.9** | | Hardware | RTX 5090 (32 GB), GPU | | Final val loss | **4.6020** (ppl 99.68) | ### ⚠️ Honest caveat: undertrained DedeProGames requested **2.6B tokens**. This checkpoint is at **50.5M tokens** — a ~50× shortfall. The GPU was occupied by other work for most of the training window, and the CPU was over-subscribed. At 4.9 tok/param, the model has learned grammar and surface fluency but not deep coherence. The val loss (4.602) is well below the 7.38 unigram floor, so it genuinely uses context — but the prose is semantically thin: wordy, repetitive, and it drifts off-topic mid-sentence. This is a **first checkpoint**, not the final deliverable. Continued training toward the 2.6B budget is planned. ## Eval (40 samples, 8 prompts × 5 seeds, temp 0.8, top-k 40) | Metric | Value | |---|---| | mean loop_frac | 0.042 | | max loop_frac | 1.0 (one sample) | | Degenerate? | **No** | | Below unigram floor? | **Yes** (4.602 < 7.38) | ### Sample outputs > "The sun is assembled by a new study because he is an associate of the study of the disease in the early years. I've been interested in having a very different study…" > "Once upon a time, she is an attack. But he is not a good idea. But it is something that does not have a moment of his own life…" > "The cat sat on the ground. The sunp is a piece of light and is not a good deal of tear. The new story of the MD's Ford…" > "def hello(): I have a lot of the best. I'm not sure what happened to me. I'll be able to do anything, but I think it's been just a lot to say…" Grammar is intact. No token loops, no broken tokens, no speaker tags. Semantically limited — expected at 4.9 tok/param. ## Usage The model uses a custom `LDT` architecture (standard LLaMA block, no special tricks). The safetensors file contains 47 tensors with tied embeddings (lm_head is not stored separately; it shares `tok.weight`). To load with a custom model class, you need a small LLaMA-style implementation matching the config above. The training script (`train_ldt10m_fixed.py`) contains the full architecture definition. ## What this is NOT - Not a 2.6B-token model (that's the target; this is the 50.5M checkpoint) - Not a general-purpose assistant (it's a raw LM, no instruction tuning) - Not a replacement for anything larger — it's a research checkpoint in a from-scratch training run