I want to be listed on "AxiomicLabs/Open_SLM_Leaderboard"
📋 Model Submission Request — GODELEV/Archaea-74M
Submitted by: GODELEV
Date: June 2, 2026
HuggingFace Model Page: https://huggingface.co/GODELEV/Archaea-74M
About the Model
Archaea-74M is a from-scratch pretrained causal language model with 74 million parameters, designed to explore the capabilities and limits of small language models (SLMs) trained entirely from the ground up — without any fine-tuning or distillation from larger models.
| Property | Value |
|---|---|
| Model ID | GODELEV/Archaea-74M |
| Architecture | Causal LM (Transformer decoder) |
| Parameters | ~74 million |
| Precision | float16 |
| Framework | Transformers (HuggingFace) |
| Training | Pretrained from scratch |
| Model Type | Base / Pretrained |
Evaluation Results
I have independently evaluated Archaea-74M using EleutherAI's lm-evaluation-harness on an NVIDIA L4 GPU (24 GB VRAM), using float16 precision and a batch size of 8. Evaluation was run on June 1, 2026. The full results JSON is attached below and is also available upon request.
Per-Task Scores
| Benchmark | Few-Shot | Metric | Score | ± Stderr |
|---|---|---|---|---|
| HellaSwag | 10-shot | acc_norm | 27.16% | ±0.44% |
| PIQA | 0-shot | acc_norm | 58.60% | ±1.15% |
| WinoGrande | 5-shot | acc | 51.14% | ±1.41% |
| BoolQ | 0-shot | acc | 56.30% | ±0.87% |
| ARC-Easy | 25-shot | acc_norm | 40.11% | ±1.01% |
| ARC-Challenge | 25-shot | acc_norm | 23.04% | ±1.23% |
| OpenBookQA | 0-shot | acc_norm | 26.00% | ±1.96% |
| CommonsenseQA | 7-shot | acc | 18.84% | ±1.12% |
| LAMBADA (OpenAI) | 0-shot | acc | 18.05% | ±0.54% |
| BLiMP | 0-shot | acc | 74.89% | ±0.14% |
| MMLU | 5-shot | acc | 25.07% | ±0.36% |
Category Averages
| Category | Average Score |
|---|---|
| Commonsense / NLI | 37.65% |
| Language Modelling | 18.05% |
| Linguistic (BLiMP) | 74.89% |
| Knowledge (MMLU) | 25.07% |
| Overall Average | 38.11% |
Note:
social_iqafailed due to a deprecated HuggingFace dataset loading script (social_i_qa.py), andarithmetic_2digitfailed as the task name has been renamed in the current version of lm-eval. Re-evaluation of these two tasks is in progress.
Evaluation Setup
Evaluation Framework : EleutherAI lm-evaluation-harness
Hardware : NVIDIA L4 GPU (24 GB VRAM) — Lighting Ai
Precision : float16
Batch Size : 8
Dataset Limit : None (full datasets)
Evaluated At : 2026-06-01T20:01:47
Total Runtime : ~36 minutes
Request
I would like to formally request the inclusion of GODELEV/Archaea-74M in the AxiomicLabs/Open_SLM_Leaderboard. The model is publicly available on HuggingFace, fully reproducible, and the evaluation results above were produced using standard lm-eval-harness tooling consistent with leaderboard methodology.
I am happy to:
- Provide the full raw
results.jsonfrom lm-eval-harness - Re-run any specific benchmarks under your preferred configuration
Thank you for maintaining this leaderboard and supporting the small language model research community. I look forward to your response.
— Akshit
Hi Akshit!
Always happy to add a new model to the leaderboard!
Benchmarks have been verified and added.
Thank You !!
No worries, excited to see what you produce next!