I want to be listed on "AxiomicLabs/Open_SLM_Leaderboard"

#2
by GODELEV - opened

📋 Model Submission Request — GODELEV/Archaea-74M

Submitted by: GODELEV
Date: June 2, 2026
HuggingFace Model Page: https://huggingface.co/GODELEV/Archaea-74M


About the Model

Archaea-74M is a from-scratch pretrained causal language model with 74 million parameters, designed to explore the capabilities and limits of small language models (SLMs) trained entirely from the ground up — without any fine-tuning or distillation from larger models.

Property Value
Model ID GODELEV/Archaea-74M
Architecture Causal LM (Transformer decoder)
Parameters ~74 million
Precision float16
Framework Transformers (HuggingFace)
Training Pretrained from scratch
Model Type Base / Pretrained

Evaluation Results

I have independently evaluated Archaea-74M using EleutherAI's lm-evaluation-harness on an NVIDIA L4 GPU (24 GB VRAM), using float16 precision and a batch size of 8. Evaluation was run on June 1, 2026. The full results JSON is attached below and is also available upon request.

Per-Task Scores

Benchmark Few-Shot Metric Score ± Stderr
HellaSwag 10-shot acc_norm 27.16% ±0.44%
PIQA 0-shot acc_norm 58.60% ±1.15%
WinoGrande 5-shot acc 51.14% ±1.41%
BoolQ 0-shot acc 56.30% ±0.87%
ARC-Easy 25-shot acc_norm 40.11% ±1.01%
ARC-Challenge 25-shot acc_norm 23.04% ±1.23%
OpenBookQA 0-shot acc_norm 26.00% ±1.96%
CommonsenseQA 7-shot acc 18.84% ±1.12%
LAMBADA (OpenAI) 0-shot acc 18.05% ±0.54%
BLiMP 0-shot acc 74.89% ±0.14%
MMLU 5-shot acc 25.07% ±0.36%

Category Averages

Category Average Score
Commonsense / NLI 37.65%
Language Modelling 18.05%
Linguistic (BLiMP) 74.89%
Knowledge (MMLU) 25.07%
Overall Average 38.11%

Note: social_iqa failed due to a deprecated HuggingFace dataset loading script (social_i_qa.py), and arithmetic_2digit failed as the task name has been renamed in the current version of lm-eval. Re-evaluation of these two tasks is in progress.


Evaluation Setup

Evaluation Framework : EleutherAI lm-evaluation-harness
Hardware             : NVIDIA L4 GPU (24 GB VRAM) — Lighting Ai
Precision            : float16
Batch Size           : 8
Dataset Limit        : None (full datasets)
Evaluated At         : 2026-06-01T20:01:47
Total Runtime        : ~36 minutes

Request

I would like to formally request the inclusion of GODELEV/Archaea-74M in the AxiomicLabs/Open_SLM_Leaderboard. The model is publicly available on HuggingFace, fully reproducible, and the evaluation results above were produced using standard lm-eval-harness tooling consistent with leaderboard methodology.

I am happy to:

  • Provide the full raw results.json from lm-eval-harness
  • Re-run any specific benchmarks under your preferred configuration

Thank you for maintaining this leaderboard and supporting the small language model research community. I look forward to your response.

— Akshit

Axiomic Labs org

Hi Akshit!
Always happy to add a new model to the leaderboard!
Benchmarks have been verified and added.

Thank You !!

Axiomic Labs org

No worries, excited to see what you produce next!

Datdanboi25 changed discussion status to closed

Sign up or log in to comment