Text Generation
Transformers
Safetensors
English
lowonmind
tiny-lm
pretrained-from-scratch
scaling-limits
custom_code
Instructions to use DedeProGames/LowOnMind-5M with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use DedeProGames/LowOnMind-5M with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="DedeProGames/LowOnMind-5M", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("DedeProGames/LowOnMind-5M", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use DedeProGames/LowOnMind-5M with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "DedeProGames/LowOnMind-5M" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DedeProGames/LowOnMind-5M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/DedeProGames/LowOnMind-5M
- SGLang
How to use DedeProGames/LowOnMind-5M with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "DedeProGames/LowOnMind-5M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DedeProGames/LowOnMind-5M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "DedeProGames/LowOnMind-5M" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "DedeProGames/LowOnMind-5M", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use DedeProGames/LowOnMind-5M with Docker Model Runner:
docker model run hf.co/DedeProGames/LowOnMind-5M
| license: apache-2.0 | |
| datasets: | |
| - HuggingFaceFW/fineweb-edu | |
| language: | |
| - en | |
| pipeline_tag: text-generation | |
| library_name: transformers | |
| tags: | |
| - lowonmind | |
| - tiny-lm | |
| - pretrained-from-scratch | |
| - scaling-limits | |
| # LowOnMind-5M | |
| A decoder-only language model with **4,920,384 parameters**, pretrained from | |
| scratch on **200M tokens** of | |
| `HuggingFaceFW/fineweb-edu` (sample-10BT). | |
| The largest model in the LowOnMind family and the third point on its scaling | |
| curve, after | |
| [LowOnMind-300k](https://huggingface.co/DedeProGames/LowOnMind-300k) and | |
| [LowOnMind-1M](https://huggingface.co/DedeProGames/LowOnMind-1M). All three | |
| share an **identical tokenizer, dataset, token budget (200M) and schedule | |
| shape**, so validation loss, bits-per-character and benchmark results are | |
| directly comparable across the series. | |
| It is also the first model in the family whose benchmark performance is | |
| statistically distinguishable from chance. | |
| ## Architecture | |
| | | 300k | 1M | 5M | | |
| |---|---:|---:|---:| | |
| | parameters | 296,960 | 985,152 | **4,920,384** | | |
| | hidden_size | 64 | 96 | 192 | | |
| | intermediate_size | 136 (2.12x) | 256 (2.67x) | 512 (2.667x) | | |
| | num_hidden_layers | 6 | 9 | 12 | | |
| | heads (q / kv) | 4 / 2 | 6 / 2 | 12 / 4 | | |
| | head_dim | 16 | 16 | 16 | | |
| | aspect ratio | 10.7 | 10.7 | **16.0** | | |
| | embedding share | 22.1% | 10.0% | **4.0%** | | |
| | vocab_size | 1024 | 1024 | 1024 (same tokenizer) | | |
| | context | 512 | 512 | 512 | | |
| | tokens seen | 200M | 200M | 200M | | |
| | tokens/param | 673 | 203 | 41 | | |
| Two deviations from the smaller siblings, both deliberate: | |
| - **Aspect ratio rises from 10.7 to 16.0.** This is the | |
| normal direction when scaling (GPT-2 small sits at 64). Holding 10.7 at this | |
| budget would require roughly 18 layers of hidden_size=160 with an | |
| implausibly wide MLP. | |
| - **intermediate/hidden is now exactly 8/3 = 2.667**, the standard SwiGLU ratio | |
| used by Llama. LowOnMind-300k was at 2.12 and LowOnMind-1M at 2.67. | |
| The **vocabulary was deliberately left at 1024** rather than raised to something | |
| more appropriate for this scale. A larger vocabulary would compress better | |
| (1024-token byte-level BPE runs about 2.35 characters per token, so 200M tokens | |
| is only ~470MB of text) and would almost certainly improve absolute results. | |
| Keeping it fixed is what makes the three-model comparison valid — the cost is | |
| that this model spends capacity assembling words from fragments that a | |
| 4096-token vocabulary would hand it for free. | |
| Modelling code is otherwise byte-identical to the two smaller siblings: GQA, | |
| SwiGLU, RMSNorm, tied embeddings, **QK-Norm** per head, **precomputed RoPE** | |
| with automatic re-expansion, residual projections initialized at | |
| `std / sqrt(2 * num_layers)`. | |
| ## Training | |
| | | | | |
| |---|---| | |
| | data | `HuggingFaceFW/fineweb-edu`, sample-10BT | | |
| | tokens | 200M (6,103 steps x 32,768) | | |
| | sequence length | 512 | | |
| | batch size | 64 | | |
| | optimizer | AdamW, betas (0.9, 0.95), wd 0.1 | | |
| | lr | 1.2e-03 peak, cosine to 1.2e-04, 250 warmup | | |
| | grad clip | 1.0 | | |
| | precision | float16 + GradScaler | | |
| | hardware | Tesla T4 | | |
| | wall clock | 27 min | | |
| At 41 tokens per parameter this run is the closest of the three to the | |
| Chinchilla-optimal ratio of roughly 20 — about 2x above it, against 10x for | |
| LowOnMind-1M and 34x for LowOnMind-300k. Train and validation loss tracked | |
| each other throughout; no overfitting. | |
| ## Results | |
| | metric | 300k | 1M | 5M | | |
| |---|---:|---:|---:| | |
| | validation loss | 3.2982 | 2.9908 | **2.5828** | | |
| | validation perplexity | 27.06 | 19.90 | **13.23** | | |
| | bits per character | 2.030 | 1.836 | **1.586** | | |
| Perplexity is not comparable across tokenizers, but it is comparable across | |
| these three models because they share one. Bits per character | |
| (loss / ln 2 / 2.35 chars-per-token) is the portable figure. | |
| Deltas: -0.4080 nats from LowOnMind-1M (5.0x the parameters), | |
| -0.7154 nats from LowOnMind-300k (16.6x). | |
| ### Real-word rate | |
| With a 1024-token byte-level vocabulary, no long word exists as a single token — | |
| the model has to assemble every one of them from fragments. The fraction of | |
| emitted words that are real English words was introduced to measure this. | |
| | | rate | | |
| |---|---:| | |
| | LowOnMind-1M | 98.0% | | |
| | LowOnMind-5M | 96.3% | | |
| | FineWeb-Edu itself (same lexicon) | 98.4% | | |
| Measured over 64 unconditional samples (5,398 words), using the same reference | |
| lexicon as LowOnMind-1M: words appearing at least 5 times in a 20k-document | |
| sample of the training corpus. | |
| **This number went down, and it should not be read as degraded spelling.** The | |
| drop is statistically real (z = 5.38, not sampling noise), but inspecting the | |
| non-words shows what happened: `illuminator` is an ordinary English word, | |
| `phillipsburg` is a US town, `shima` is a common element of Japanese place | |
| names. They are counted as errors only because they fall below the reference | |
| lexicon's frequency-5 threshold. The remainder (`hymenola`, `almanine`, | |
| `perleti`, `amiravicis`) skew toward proper-noun and Latinate-technical | |
| morphology rather than the malformed common words the metric was built to catch | |
| — LowOnMind-300k produced things like `landship` and `parsetic`, failures of a | |
| different kind. | |
| **The metric has a floor problem as well as a ceiling problem.** As a model | |
| improves it emits rarer real vocabulary — names, places, technical terms — which | |
| a frequency-thresholded lexicon scores as wrong. So the measured rate can fall | |
| while actual quality rises. Comparing against a full dictionary with proper-noun | |
| handling, rather than a corpus-frequency cutoff, would be the fix. The 96.3% | |
| figure is reported as-measured for continuity, but it should not be used to rank | |
| these models. | |
| ## BananaMind Base Bench 1.1 | |
| Evaluated on [BananaMind/BananaMind-Base-Bench-1.1](https://huggingface.co/datasets/BananaMind/BananaMind-Base-Bench-1.1), | |
| the same 350-item English continuation-likelihood benchmark used across the | |
| family, with identical scoring: context and each of the four continuations | |
| tokenized separately with `add_special_tokens=False`, no BOS, selection by | |
| highest mean conditional token log-probability. | |
| Run validity: dataset SHA-256 matched, full schema validation passed, no context | |
| required truncation against the 512-token window. | |
| | Category | 300k | 1M | 5M | z vs chance (5M) | Elo (5M) | | |
| |---|---:|---:|---:|---:|---:| | |
| | language_completion | 46.0% | 52.0% | **62.0%** | **+6.04** | **1008** | | |
| | world_knowledge | 22.0% | 22.0% | **38.0%** | +2.12 | 881 | | |
| | context_tracking | 14.0% | 24.0% | 32.0% | +1.14 | 851 | | |
| | quantitative | 32.0% | 28.0% | 28.0% | +0.49 | 872 | | |
| | logical_reasoning | 24.0% | 28.0% | 26.0% | +0.16 | 900 | | |
| | commonsense | 34.0% | 28.0% | 24.0% | -0.16 | 758 | | |
| | code_completion | 14.0% | 20.0% | 16.0% | -1.47 | 805 | | |
| | | 300k | 1M | 5M | | |
| |---|---:|---:|---:| | |
| | Overall Elo | 833 | 843 | **863** | | |
| | Chance-level Elo (this grid) | 805 | 805 | 805 | | |
| | Raw accuracy | 26.6% | 28.9% | **32.3%** | | |
| | 95% CI | [22.0, 31.2] | [24.2, 33.6] | **[27.4, 37.2]** | | |
| | z vs. chance | +0.69 | +1.68 | **+3.15** | | |
| | significant vs. chance | no | no | **yes** | | |
| Difficulty split: easy 30.8%, medium 33.3%, hard 32.8%. | |
| ## Usage | |
| ```python | |
| from transformers import AutoModelForCausalLM, AutoTokenizer | |
| tok = AutoTokenizer.from_pretrained("DedeProGames/LowOnMind-5M") | |
| model = AutoModelForCausalLM.from_pretrained("DedeProGames/LowOnMind-5M", trust_remote_code=True) | |
| ids = tok("The ", return_tensors="pt").input_ids | |
| print(tok.decode(model.generate(ids, max_new_tokens=64, use_cache=False)[0])) | |
| ``` | |
| `trust_remote_code=True` is required — the architecture ships as custom modeling | |
| code in the repository. `use_cache=False` is required: this implementation has no | |
| KV cache and recomputes the full window at each generation step. | |
| ## Limitations | |
| At ~5M parameters this is still a research artifact, not a usable model. Expect | |
| fluent local syntax and register-appropriate structure, but **no reliable | |
| coherence across a paragraph**, no dependable factual knowledge, and no ability | |
| to track state across a passage. Benchmark accuracy of 32.3% is above chance and | |
| far below usefulness. The 1024-token vocabulary caps absolute quality below what | |
| this parameter count could otherwise reach. | |
| The 512-token context and absent KV cache also make it unsuitable for any real | |
| workload. |