Small language model evaluation

Research index

Compact
Language Model
Index

A fixed evaluation of compact language models across six benchmarks, with chance-normalized scores and 95% intervals

Explore the leaderboard
Evaluation breadth
6core benchmarks
Linguistic depth
12BLiMP categories
Current sample
20models measured

Ranked results

Core leaderboard

RankTrack
1–2LFM2 350M LiquidAI/LFM2-350M354Minstruction43.242.0–44.232.030.6–33.339.235.0–43.255.252.6–57.719.515.7–23.331.327.9–34.958.157.2–58.4
1–2Qwen2.5 0.5B Qwen/Qwen2.5-0.5B494Mbase43.041.9–44.036.234.8–37.538.834.7–42.844.641.9–47.29.35.8–13.126.222.9–29.768.767.9–68.9
3Gemma 3 270M google/gemma-3-270m270Mbase38.937.8–39.921.920.7–23.337.032.9–41.243.040.4–45.64.00.6–7.328.224.8–31.764.363.5–64.5
4SmolLM2 135M HuggingFaceTB/SmolLM2-135M135Mbase37.836.7–38.824.323.0–25.636.332.1–40.644.942.3–47.56.32.6–9.719.416.2–22.862.761.9–63.1
5Limen0.2B UniversalComputingResearch/Limen0.2B223Mbase36.635.5–37.622.621.3–23.934.730.4–38.837.835.2–40.46.32.7–9.917.714.3–21.166.665.8–66.9
6GPT-X2 125M AxiomicLabs/GPT-X2-125M125Mbase35.534.3–36.520.719.4–21.934.229.7–38.435.532.8–38.23.4-0.1–7.118.314.9–21.665.865.0–66.1
7GPT-2 Medium 355M openai-community/gpt2-medium355Mbase33.432.3–34.419.117.9–20.333.228.9–37.524.622.1–27.30.0-3.1–3.314.611.5–18.070.469.8–70.8
8MobileLLM-R1 140M Base facebook/MobileLLM-R1-140M-base140Mbase31.230.1–32.212.110.9–13.326.622.2–30.833.230.5–35.7-1.1-4.6–2.319.316.0–22.861.660.7–61.8
9Baguettotron PleIAs/Baguettotron321Minstruction29.628.4–30.713.912.6–15.124.219.6–28.734.131.5–36.77.23.8–10.613.310.0–16.656.956.0–57.2
10–11GPT-2 124M openai-community/gpt2124Mbase27.926.8–28.98.27.0–9.425.020.7–29.419.316.7–21.9-3.4-6.6–-0.212.39.2–15.567.566.8–67.8
10–11Supra 50M Base SupraLabs/Supra-50M-Base51.8Mbase27.326.2–28.38.87.6–10.023.418.9–28.027.524.7–30.2-0.1-3.5–3.212.59.2–15.759.057.9–59.1
12Veyra2 Apricot 50M Base veyra-ai/Veyra2-Apricot-50M-Base49.3Mbase26.124.9–27.18.47.2–9.523.419.0–27.923.720.9–26.4-2.0-5.2–1.110.87.7–14.058.757.7–58.9
13–15Falcon H1 Tiny R 90M tiiuae/Falcon-H1-Tiny-R-90M90.0Minstruction20.819.6–21.87.76.6–8.921.517.0–25.919.016.3–21.5-0.1-3.4–3.210.47.4–13.642.541.6–43.0
13–15SLM 10M liodon-ai/slm-10m9.97Mbase20.319.1–21.23.12.0–4.214.19.7–18.514.311.6–16.8-1.9-5.1–1.37.54.5–10.654.453.5–54.7
13–16Pythia 160M EleutherAI/pythia-160m160Mbase20.219.1–21.27.36.2–8.516.912.5–21.215.813.2–18.4-2.6-5.7–0.611.78.4–14.845.744.8–46.2
15–16GPT-S2 5M AxiomicLabs/GPT-S2-5M5.38Mbase19.318.2–20.33.42.3–4.512.98.3–17.411.38.6–13.9-3.8-6.9–-0.76.03.2–9.254.653.9–55.0
17–18nanowhale 100M Base HuggingFaceTB/nanowhale-100m-base110Mbase17.916.8–18.92.81.6–4.014.710.2–19.316.413.7–19.0-3.4-6.5–-0.29.76.6–12.941.841.0–42.3
17–18Pythia 31M EleutherAI/pythia-31m31.0Mbase16.915.8–18.02.91.8–4.113.69.1–18.112.19.7–14.5-4.4-7.6–-1.18.15.1–11.143.342.5–43.9
19Atom 3.4M UniversalComputingResearch/Atom3.4m3.41Mbase14.313.3–15.33.52.4–4.511.46.9–15.910.88.4–13.2-4.2-7.4–-1.15.92.9–9.236.435.6–37.0
20GPT-S 1.4M AxiomicLabs/GPT-S-1.4M1.43Mbase12.311.2–13.32.51.3–3.610.35.7–14.79.06.5–11.5-4.0-7.1–-0.74.31.5–7.432.031.2–32.6

ARC-E* = ARC-Easy · ARC-C* = ARC-Challenge · CQA* = CommonsenseQA

Chart metric

Select a benchmark to update both the scaling chart and score comparison below.

Model scaling

Parameters and Average score

Base Instruction-tuned Pareto trend
Chance-normalized score-5.08.021.034.047.01.00M66.2M233M502MParameters (square-root scale)

Model score comparison

The top 10 matching models are shown by default. This comparison uses the shared metric selected above.

1 LFM2 350M
43.2
2 Qwen2.5 0.5B
43.0
3 Gemma 3 270M
38.9
4 SmolLM2 135M
37.8
5 Limen0.2B
36.6
6 GPT-X2 125M
35.5
7 GPT-2 Medium 355M
33.4
8 MobileLLM-R1 140M Base
31.2
9 Baguettotron
29.6
10 GPT-2 124M
27.9

Scoring method

Chance-normalized composite score

100 × (accuracy − chance) / (1 − chance)

For each benchmark, the observed result is normalized against its chance level: chance maps to 0 and a perfect result maps to 100.

HellaSwag, PIQA, and CommonsenseQA form a commonsense domain worth 50% of the composite, with each benchmark weighted equally within that domain. Science and world knowledge contribute 25% after combining ARC-Easy (75%) and ARC-Challenge (25%). BLiMP contributes the remaining 25% for grammatical competence and is macro-averaged across its 12 linguistic categories.

CommonsenseQA scores the likelihood of each full answer text with length normalization, rather than scoring the A–E answer label. This prevents a model’s letter preference from being mistaken for commonsense capability.

Each benchmark interval is a 95% percentile bootstrap. Model comparisons use paired bootstrap resampling on matched evaluation items. Results below chance remain negative, and rank ranges reflect comparisons that do not distinguish neighbouring models.

Evaluation protocol

Curated evaluation

CompactLMIndex is a curated benchmark, not an open intake list. We are establishing a Pareto-frontier baseline across model sizes and highlighting training or inference strategies that put a model ahead of its peers at a comparable parameter count.

  • Complete core suite required for a rank
  • Zero-shot, non-chat log-likelihood prompts
  • Paired bootstrap comparisons on matched questions
  • Benchmark-specific training or fine-tuning is discouraged and may exclude a model from comparison

Every published result is independently checked through our internal verification process. Final inclusion and presentation decisions remain with Universal Computing Research.

If you have interesting results for a model you think is suitable for this benchmark, you can open a discussion on the Space for consideration.

Credits & scope

Efficiency at smaller scales

CompactLMIndex looks at smaller language models. At this scale, architecture, training, and inference choices can make a real difference. Every model follows the same evaluation protocol, so results are compared fairly rather than by size alone.

Inspired by the Open LLM Leaderboard, Axiomic Labs Open SLM Leaderboard, and OpenAI Parameter Golf.