How to use from the
Use from the
Transformers library
# Use a pipeline as a high-level helper
from transformers import pipeline

pipe = pipeline("text-generation", model="User01110/100M-exp", trust_remote_code=True)
# pip install -U transformers accelerate
# Load model directly
from transformers import AutoModelForCausalLM
model = AutoModelForCausalLM.from_pretrained("User01110/100M-exp", trust_remote_code=True, device_map="auto")
Quick Links

100M-exp

A 98M-parameter base language model, trained from scratch on 21B tokens.

100M-exp is an experiment in small-model design: a hybrid sliding-window / full-attention transformer with a few additions meant to help at this scale. It was trained for 40k steps as a branch of a longer run, ending with a learning-rate cooldown. It scores 22.47 on the Open SLM Leaderboard's Intelligence Index (self-evaluated).


At a glance

Parameters 98.2M (tied input/output embedding)
Layers 20 (18 sliding-window + 2 full attention)
Width 576 hidden, 9 query heads / 3 key-value heads (head dim 64)
Context 8,192 tokens
Vocabulary 32,000 (Mistral BPE)
Training tokens ~21B (40,000 steps x 524,288 tokens)
Optimizer Muon (matrices) + AdamW (everything else)
Type Base model (not instruction-tuned)

Benchmarks

All scores are zero-shot, acc_norm, measured with our evaluation script that mirrors the leaderboard's Intelligence Index.

Benchmark Score Random chance What it tests
HellaSwag 37.78 25 Common-sense sentence completion
ARC-Easy 49.83 25 Grade-school science questions
ARC-Challenge 27.22 25 Harder science questions
PIQA 66.97 50 Physical common sense
ArithMark-3 40.00 25 Arithmetic reasoning
WikiText (bits/byte, lower is better) 0.9226 Raw language modeling

Intelligence Index: 22.47

Self-evaluated; the leaderboard's own run may differ by a few tenths.

How the Index is built

The Index measures how far above random guessing a model is, as a share of the gap between guessing and perfect, averaged across tasks:

How the Index is built: HellaSwag 17.0, ARC 18.0, PIQA 33.9 and ArithMark 20.0 (weighted 0.65) points above chance, averaged to 22.47

In other words, 100M-exp covers about 22% of the distance between random guessing and a perfect score.


More benchmarks

Beyond the leaderboard tasks: zero-shot, lm-evaluation-harness, same model.

Benchmark Metric Score Random chance What it tests
SciQ acc / acc_norm 83.4 / 76.5 25 Science questions with a supporting passage
BLiMP (67 subtasks) acc 81.3 50 Grammar: picking the grammatical sentence of a minimal pair
LAMBADA (OpenAI) acc 32.8 ~0 Predicting a passage's final word from long context
LAMBADA (OpenAI) perplexity (lower is better) 37.1
BoolQ acc 57.6 62 (always "yes") Yes/no reading comprehension

Architecture

100M-exp is a pre-norm decoder-only transformer with a hybrid attention layout: most layers only look at a local window, and two layers see everything.

100M-exp layer layout: 18 sliding-window blocks and 2 full-attention blocks (7 and 15), with tied embedding and output head

Inside every block:

Inside a block: RMSNorm, fused qkv, causal conv, HoPE, attention with a learned sink and value residual, then RMSNorm and SwiGLU MLP, each with a residual connection

What each piece does, in plain terms:

  • Sliding-window + full attention. 18 of 20 layers attend only to the previous 512 tokens, which is cheap and good at local grammar and phrasing. Two full-attention layers (7 and 15) connect the whole 8k context, so long-range information still flows.
  • Attention sinks. Each head has a learned "nowhere" slot it can attend to, so it doesn't have to dump attention on real tokens when nothing is relevant (gpt-oss style).
  • HoPE positions. A RoPE variant: rotary dimensions slower than one turn per 8k tokens are left un-rotated, giving the model some position-free channels for pure content matching.
  • Short causal convolution. A tiny 4-token convolution on queries, keys and values lets each token blend in its immediate neighbours before attention, which tends to help small models.
  • Value residual (ResFormer). Every block can mix in the first block's values, keeping early token information easy to reach in deep layers.
  • Document masking. Documents packed into the same 8k sequence can never attend to each other, so the model never learns false connections between unrelated texts.
  • GQA 9:3. Three key-value heads shared across nine query heads, which keeps attention lean.

Training

Data mix

Training data mix: ClimbMix 45%, DCLM-Edu 15%, Nemotron-CC high quality 15%, NExtLong 15%, CC-Math-Finest 10%

A broad, quality-filtered web mix, an educational slice, long documents for the 8k context, and a math slice. Every training batch draws from every source.

Recipe

Batch 524,288 tokens per step (8,192 context)
Steps 40,000 (~21B tokens, ~214 tokens per parameter)
Optimizer Muon (momentum 0.95, weight decay 0.1) for attention and MLP matrices; AdamW (betas 0.9/0.95, no decay) for embeddings, norms, sinks and convs
Peak LR 1e-3
Precision bf16 with fp32 loss accumulation
Gradient clipping 1.0

Schedule

Learning-rate schedule: 500-step warmup, constant 1e-3 to 30k, then a 1 - sqrt cooldown to zero at 40k; Index 20.61 before the cooldown and 22.47 after

100M-exp is a cooldown branch: it starts from the step-30k checkpoint of a longer 100k-step run and decays the learning rate to zero over 10k steps. Most of the final gain came from the cooldown:

Checkpoint Intelligence Index WikiText bits/byte
30k, before cooldown (best weight average) 20.61 0.9443
32k, cooldown starting 20.87 0.9510
39k, near the end 21.94 0.9241
40k, final (100M-exp) 22.47 0.9226

Usage

The model uses custom code, so trust_remote_code=True is required.

Requirements: PyTorch 2.5 or newer (attention runs on FlexAttention) and a CUDA GPU; tested with transformers 5.14. Generation uses a KV cache. Batches must not be padded: generate one prompt at a time, or batch only prompts of equal length.

import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "User01110/100M-exp"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, trust_remote_code=True, dtype=torch.bfloat16).cuda().eval()

prompt = "The process of photosynthesis"
ids = tok(prompt, return_tensors="pt").input_ids.cuda()
out = model.generate(ids, max_new_tokens=60, do_sample=True, temperature=0.7, top_p=0.9)
print(tok.decode(out[0], skip_special_tokens=True))

Limitations

  • A base model. It continues text; it was not trained to follow instructions or chat.
  • Small. At 98M parameters it makes factual mistakes, loses the thread in long generations, and is weak at multi-step reasoning. Treat it as a research artifact, not a source of truth.
  • Mostly English, reflecting its training data.
  • Inherits web-data biases. The training mix is filtered web text, so outputs can reflect the biases and errors in that data.

Acknowledgements

100M-exp combines published ideas; none of the components below are new here.

Architecture

Training

Data and tokenizer

Evaluation


License

Apache-2.0 for the weights, code and model card. The tokenizer is Mistral's, also Apache-2.0. Each training dataset keeps its own license.

Downloads last month
-
Safetensors
Model size
98.2M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train User01110/100M-exp

Papers for User01110/100M-exp