--- license: apache-2.0 datasets: - HuggingFaceFW/fineweb-edu language: - en pipeline_tag: text-generation library_name: transformers tags: - lowonmind - tiny-lm - pretrained-from-scratch - scaling-limits --- # LowOnMind-5M A decoder-only language model with **4,920,384 parameters**, pretrained from scratch on **200M tokens** of `HuggingFaceFW/fineweb-edu` (sample-10BT). The largest model in the LowOnMind family and the third point on its scaling curve, after [LowOnMind-300k](https://huggingface.co/DedeProGames/LowOnMind-300k) and [LowOnMind-1M](https://huggingface.co/DedeProGames/LowOnMind-1M). All three share an **identical tokenizer, dataset, token budget (200M) and schedule shape**, so validation loss, bits-per-character and benchmark results are directly comparable across the series. It is also the first model in the family whose benchmark performance is statistically distinguishable from chance. ## Architecture | | 300k | 1M | 5M | |---|---:|---:|---:| | parameters | 296,960 | 985,152 | **4,920,384** | | hidden_size | 64 | 96 | 192 | | intermediate_size | 136 (2.12x) | 256 (2.67x) | 512 (2.667x) | | num_hidden_layers | 6 | 9 | 12 | | heads (q / kv) | 4 / 2 | 6 / 2 | 12 / 4 | | head_dim | 16 | 16 | 16 | | aspect ratio | 10.7 | 10.7 | **16.0** | | embedding share | 22.1% | 10.0% | **4.0%** | | vocab_size | 1024 | 1024 | 1024 (same tokenizer) | | context | 512 | 512 | 512 | | tokens seen | 200M | 200M | 200M | | tokens/param | 673 | 203 | 41 | Two deviations from the smaller siblings, both deliberate: - **Aspect ratio rises from 10.7 to 16.0.** This is the normal direction when scaling (GPT-2 small sits at 64). Holding 10.7 at this budget would require roughly 18 layers of hidden_size=160 with an implausibly wide MLP. - **intermediate/hidden is now exactly 8/3 = 2.667**, the standard SwiGLU ratio used by Llama. LowOnMind-300k was at 2.12 and LowOnMind-1M at 2.67. The **vocabulary was deliberately left at 1024** rather than raised to something more appropriate for this scale. A larger vocabulary would compress better (1024-token byte-level BPE runs about 2.35 characters per token, so 200M tokens is only ~470MB of text) and would almost certainly improve absolute results. Keeping it fixed is what makes the three-model comparison valid — the cost is that this model spends capacity assembling words from fragments that a 4096-token vocabulary would hand it for free. Modelling code is otherwise byte-identical to the two smaller siblings: GQA, SwiGLU, RMSNorm, tied embeddings, **QK-Norm** per head, **precomputed RoPE** with automatic re-expansion, residual projections initialized at `std / sqrt(2 * num_layers)`. ## Training | | | |---|---| | data | `HuggingFaceFW/fineweb-edu`, sample-10BT | | tokens | 200M (6,103 steps x 32,768) | | sequence length | 512 | | batch size | 64 | | optimizer | AdamW, betas (0.9, 0.95), wd 0.1 | | lr | 1.2e-03 peak, cosine to 1.2e-04, 250 warmup | | grad clip | 1.0 | | precision | float16 + GradScaler | | hardware | Tesla T4 | | wall clock | 27 min | At 41 tokens per parameter this run is the closest of the three to the Chinchilla-optimal ratio of roughly 20 — about 2x above it, against 10x for LowOnMind-1M and 34x for LowOnMind-300k. Train and validation loss tracked each other throughout; no overfitting. ## Results | metric | 300k | 1M | 5M | |---|---:|---:|---:| | validation loss | 3.2982 | 2.9908 | **2.5828** | | validation perplexity | 27.06 | 19.90 | **13.23** | | bits per character | 2.030 | 1.836 | **1.586** | Perplexity is not comparable across tokenizers, but it is comparable across these three models because they share one. Bits per character (loss / ln 2 / 2.35 chars-per-token) is the portable figure. Deltas: -0.4080 nats from LowOnMind-1M (5.0x the parameters), -0.7154 nats from LowOnMind-300k (16.6x). ### Real-word rate With a 1024-token byte-level vocabulary, no long word exists as a single token — the model has to assemble every one of them from fragments. The fraction of emitted words that are real English words was introduced to measure this. | | rate | |---|---:| | LowOnMind-1M | 98.0% | | LowOnMind-5M | 96.3% | | FineWeb-Edu itself (same lexicon) | 98.4% | Measured over 64 unconditional samples (5,398 words), using the same reference lexicon as LowOnMind-1M: words appearing at least 5 times in a 20k-document sample of the training corpus. **This number went down, and it should not be read as degraded spelling.** The drop is statistically real (z = 5.38, not sampling noise), but inspecting the non-words shows what happened: `illuminator` is an ordinary English word, `phillipsburg` is a US town, `shima` is a common element of Japanese place names. They are counted as errors only because they fall below the reference lexicon's frequency-5 threshold. The remainder (`hymenola`, `almanine`, `perleti`, `amiravicis`) skew toward proper-noun and Latinate-technical morphology rather than the malformed common words the metric was built to catch — LowOnMind-300k produced things like `landship` and `parsetic`, failures of a different kind. **The metric has a floor problem as well as a ceiling problem.** As a model improves it emits rarer real vocabulary — names, places, technical terms — which a frequency-thresholded lexicon scores as wrong. So the measured rate can fall while actual quality rises. Comparing against a full dictionary with proper-noun handling, rather than a corpus-frequency cutoff, would be the fix. The 96.3% figure is reported as-measured for continuity, but it should not be used to rank these models. ## BananaMind Base Bench 1.1 Evaluated on [BananaMind/BananaMind-Base-Bench-1.1](https://huggingface.co/datasets/BananaMind/BananaMind-Base-Bench-1.1), the same 350-item English continuation-likelihood benchmark used across the family, with identical scoring: context and each of the four continuations tokenized separately with `add_special_tokens=False`, no BOS, selection by highest mean conditional token log-probability. Run validity: dataset SHA-256 matched, full schema validation passed, no context required truncation against the 512-token window. | Category | 300k | 1M | 5M | z vs chance (5M) | Elo (5M) | |---|---:|---:|---:|---:|---:| | language_completion | 46.0% | 52.0% | **62.0%** | **+6.04** | **1008** | | world_knowledge | 22.0% | 22.0% | **38.0%** | +2.12 | 881 | | context_tracking | 14.0% | 24.0% | 32.0% | +1.14 | 851 | | quantitative | 32.0% | 28.0% | 28.0% | +0.49 | 872 | | logical_reasoning | 24.0% | 28.0% | 26.0% | +0.16 | 900 | | commonsense | 34.0% | 28.0% | 24.0% | -0.16 | 758 | | code_completion | 14.0% | 20.0% | 16.0% | -1.47 | 805 | | | 300k | 1M | 5M | |---|---:|---:|---:| | Overall Elo | 833 | 843 | **863** | | Chance-level Elo (this grid) | 805 | 805 | 805 | | Raw accuracy | 26.6% | 28.9% | **32.3%** | | 95% CI | [22.0, 31.2] | [24.2, 33.6] | **[27.4, 37.2]** | | z vs. chance | +0.69 | +1.68 | **+3.15** | | significant vs. chance | no | no | **yes** | Difficulty split: easy 30.8%, medium 33.3%, hard 32.8%. ## Usage ```python from transformers import AutoModelForCausalLM, AutoTokenizer tok = AutoTokenizer.from_pretrained("DedeProGames/LowOnMind-5M") model = AutoModelForCausalLM.from_pretrained("DedeProGames/LowOnMind-5M", trust_remote_code=True) ids = tok("The ", return_tensors="pt").input_ids print(tok.decode(model.generate(ids, max_new_tokens=64, use_cache=False)[0])) ``` `trust_remote_code=True` is required — the architecture ships as custom modeling code in the repository. `use_cache=False` is required: this implementation has no KV cache and recomputes the full window at each generation step. ## Limitations At ~5M parameters this is still a research artifact, not a usable model. Expect fluent local syntax and register-appropriate structure, but **no reliable coherence across a paragraph**, no dependable factual knowledge, and no ability to track state across a passage. Benchmark accuracy of 32.3% is above chance and far below usefulness. The 1024-token vocabulary caps absolute quality below what this parameter count could otherwise reach. The 512-token context and absent KV cache also make it unsuitable for any real workload.