Instructions to use User01110/100M-exp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use User01110/100M-exp with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="User01110/100M-exp", trust_remote_code=True)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("User01110/100M-exp", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use User01110/100M-exp with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "User01110/100M-exp" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/100M-exp", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/User01110/100M-exp
- SGLang
How to use User01110/100M-exp with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "User01110/100M-exp" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/100M-exp", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "User01110/100M-exp" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/100M-exp", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use User01110/100M-exp with Docker Model Runner:
docker model run hf.co/User01110/100M-exp
100M-exp
A 98M-parameter base language model, trained from scratch on 21B tokens.
100M-exp is an experiment in small-model design: a hybrid sliding-window / full-attention transformer with a few additions meant to help at this scale. It was trained for 40k steps as a branch of a longer run, ending with a learning-rate cooldown. It scores 22.47 on the Open SLM Leaderboard's Intelligence Index (self-evaluated).
At a glance
| Parameters | 98.2M (tied input/output embedding) |
| Layers | 20 (18 sliding-window + 2 full attention) |
| Width | 576 hidden, 9 query heads / 3 key-value heads (head dim 64) |
| Context | 8,192 tokens |
| Vocabulary | 32,000 (Mistral BPE) |
| Training tokens | ~21B (40,000 steps x 524,288 tokens) |
| Optimizer | Muon (matrices) + AdamW (everything else) |
| Type | Base model (not instruction-tuned) |
Benchmarks
All scores are zero-shot, acc_norm, measured with our evaluation script that mirrors the leaderboard's Intelligence Index.
| Benchmark | Score | Random chance | What it tests |
|---|---|---|---|
| HellaSwag | 37.78 | 25 | Common-sense sentence completion |
| ARC-Easy | 49.83 | 25 | Grade-school science questions |
| ARC-Challenge | 27.22 | 25 | Harder science questions |
| PIQA | 66.97 | 50 | Physical common sense |
| ArithMark-3 | 40.00 | 25 | Arithmetic reasoning |
| WikiText (bits/byte, lower is better) | 0.9226 | Raw language modeling |
Intelligence Index: 22.47
Self-evaluated; the leaderboard's own run may differ by a few tenths.
How the Index is built
The Index measures how far above random guessing a model is, as a share of the gap between guessing and perfect, averaged across tasks:
In other words, 100M-exp covers about 22% of the distance between random guessing and a perfect score.
More benchmarks
Beyond the leaderboard tasks: zero-shot, lm-evaluation-harness, same model.
| Benchmark | Metric | Score | Random chance | What it tests |
|---|---|---|---|---|
| SciQ | acc / acc_norm | 83.4 / 76.5 | 25 | Science questions with a supporting passage |
| BLiMP (67 subtasks) | acc | 81.3 | 50 | Grammar: picking the grammatical sentence of a minimal pair |
| LAMBADA (OpenAI) | acc | 32.8 | ~0 | Predicting a passage's final word from long context |
| LAMBADA (OpenAI) | perplexity (lower is better) | 37.1 | ||
| BoolQ | acc | 57.6 | 62 (always "yes") | Yes/no reading comprehension |
Architecture
100M-exp is a pre-norm decoder-only transformer with a hybrid attention layout: most layers only look at a local window, and two layers see everything.
Inside every block:
What each piece does, in plain terms:
- Sliding-window + full attention. 18 of 20 layers attend only to the previous 512 tokens, which is cheap and good at local grammar and phrasing. Two full-attention layers (7 and 15) connect the whole 8k context, so long-range information still flows.
- Attention sinks. Each head has a learned "nowhere" slot it can attend to, so it doesn't have to dump attention on real tokens when nothing is relevant (gpt-oss style).
- HoPE positions. A RoPE variant: rotary dimensions slower than one turn per 8k tokens are left un-rotated, giving the model some position-free channels for pure content matching.
- Short causal convolution. A tiny 4-token convolution on queries, keys and values lets each token blend in its immediate neighbours before attention, which tends to help small models.
- Value residual (ResFormer). Every block can mix in the first block's values, keeping early token information easy to reach in deep layers.
- Document masking. Documents packed into the same 8k sequence can never attend to each other, so the model never learns false connections between unrelated texts.
- GQA 9:3. Three key-value heads shared across nine query heads, which keeps attention lean.
Training
Data mix
A broad, quality-filtered web mix, an educational slice, long documents for the 8k context, and a math slice. Every training batch draws from every source.
Recipe
| Batch | 524,288 tokens per step (8,192 context) |
| Steps | 40,000 (~21B tokens, ~214 tokens per parameter) |
| Optimizer | Muon (momentum 0.95, weight decay 0.1) for attention and MLP matrices; AdamW (betas 0.9/0.95, no decay) for embeddings, norms, sinks and convs |
| Peak LR | 1e-3 |
| Precision | bf16 with fp32 loss accumulation |
| Gradient clipping | 1.0 |
Schedule
100M-exp is a cooldown branch: it starts from the step-30k checkpoint of a longer 100k-step run and decays the learning rate to zero over 10k steps. Most of the final gain came from the cooldown:
| Checkpoint | Intelligence Index | WikiText bits/byte |
|---|---|---|
| 30k, before cooldown (best weight average) | 20.61 | 0.9443 |
| 32k, cooldown starting | 20.87 | 0.9510 |
| 39k, near the end | 21.94 | 0.9241 |
40k, final (100M-exp) |
22.47 | 0.9226 |
Usage
The model uses custom code, so trust_remote_code=True is required.
Requirements: PyTorch 2.5 or newer (attention runs on FlexAttention) and a CUDA GPU; tested with transformers 5.14. Generation uses a KV cache. Batches must not be padded: generate one prompt at a time, or batch only prompts of equal length.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "User01110/100M-exp"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True, dtype=torch.bfloat16).cuda().eval()
prompt = "The process of photosynthesis"
ids = tok(prompt, return_tensors="pt").input_ids.cuda()
out = model.generate(ids, max_new_tokens=60, do_sample=True, temperature=0.7, top_p=0.9)
print(tok.decode(out[0], skip_special_tokens=True))
Limitations
- A base model. It continues text; it was not trained to follow instructions or chat.
- Small. At 98M parameters it makes factual mistakes, loses the thread in long generations, and is weak at multi-step reasoning. Treat it as a research artifact, not a source of truth.
- Mostly English, reflecting its training data.
- Inherits web-data biases. The training mix is filtered web text, so outputs can reflect the biases and errors in that data.
Acknowledgements
100M-exp combines published ideas; none of the components below are new here.
Architecture
- Sliding-window and local/global attention: Longformer, Mistral 7B
- Attention sinks: Efficient Streaming Language Models with Attention Sinks; learned per-head sinks as in gpt-oss
- Rotary positions: RoFormer; HoPE
- Short convolution on queries/keys/values: Primer
- Value residual: Value Residual Learning (ResFormer)
- Grouped-query attention: GQA
- SwiGLU: GLU Variants Improve Transformer; RMSNorm: Root Mean Square Layer Normalization
Training
- Muon optimizer: Keller Jordan's Muon; Muon is Scalable for LLM Training
- Warmup-stable-decay schedule and 1 - sqrt cooldown: Scaling Laws and Compute-Optimal Training Beyond Fixed Training Durations
- PyTorch FlexAttention; Liger Kernel fused cross-entropy
Data and tokenizer
- Datasets: ClimbMix, DCLM-Edu, Nemotron-CC, NExtLong, CC-Math-Finest
- Tokenizer: Mistral-7B-v0.1 (Apache-2.0)
Evaluation
- lm-evaluation-harness; the Open SLM Leaderboard's Intelligence Index
License
Apache-2.0 for the weights, code and model card. The tokenizer is Mistral's, also Apache-2.0. Each training dataset keeps its own license.
- Downloads last month
- -




