Submission: ThingAI Quark Family โ€” Quark-50m & Quark-135m

#6
by ThingsAI - opened

Submission โ€” ThingAI Quark Model Family

We would like to submit two models from the Quark family for inclusion in the Open SLM Leaderboard. Both models are fully open-weights, trained from scratch on sovereign hardware, and designed for resource-constrained deployment scenarios.

Organization

Shared Architecture

Both models share the same core architecture: Grouped Query Attention (GQA), SwiGLU activations, RMSNorm, Rotary Position Embeddings (RoPE), and weight tying between the embedding and language model head. Trained in BF16 precision with a 2048-token context window.


Model 1: ThingAI/Quark-50m

Category Benchmark Metric Score
Linguistics & Grammar BLiMP Accuracy 68.12%
Commonsense & Reasoning PIQA Norm. Accuracy 57.83%
Commonsense & Reasoning COPA Accuracy 57.00%
Commonsense & Reasoning BoolQ Accuracy 52.17%
Commonsense & Reasoning WinoGrande Accuracy 47.36%
Commonsense & Reasoning HellaSwag Norm. Accuracy 28.49%
Commonsense & Reasoning RACE Accuracy 26.41%
Commonsense & Reasoning CommonsenseQA Accuracy 20.31%
Academic & Knowledge SciQ Norm. Accuracy 49.00%
Academic & Knowledge ARC-Easy Norm. Accuracy 36.49%
Academic & Knowledge MMLU Accuracy 25.64%
Academic & Knowledge ARC-Challenge Norm. Accuracy 25.17%
Academic & Knowledge OpenBookQA Norm. Accuracy 25.40%
Language Modeling LAMBADA Accuracy 15.87%
Language Modeling WikiText-2 Word Perplexity 251.76

Quark-50m evaluation performed by @GODELEV using the standard EleutherAI lm-evaluation-harness.


Model 2: ThingAI/Quark-135m

Category Benchmark Metric Score
Commonsense & Reasoning PIQA Norm. Accuracy 61.26%
Commonsense & Reasoning WinoGrande Accuracy 50.20%
Commonsense & Reasoning HellaSwag Norm. Accuracy 31.37%
Commonsense & Reasoning CommonsenseQA Accuracy 20.56%
Academic & Knowledge ARC-Easy Norm. Accuracy 41.46%
Academic & Knowledge ARC-Challenge Norm. Accuracy 25.09%
Academic & Knowledge OpenBookQA Norm. Accuracy 27.20%
Academic & Knowledge MMLU (avg) Accuracy 23.17%
Academic & Knowledge MMLU Humanities Accuracy 24.23%
Academic & Knowledge MMLU Social Sciences Accuracy 22.59%
Academic & Knowledge MMLU STEM Accuracy 22.04%
Academic & Knowledge MMLU Other Accuracy 23.27%
Knowledge Retrieval TriviaQA Exact Match 0.07%

Quark-135m evaluation performed internally using lm-evaluation-harness.


Scaling Comparison (50M โ†’ 135M)

Benchmark Quark-50m Quark-135m ฮ”
PIQA 57.83% 61.26% +3.43
HellaSwag 28.49% 31.37% +2.88
ARC-Easy 36.49% 41.46% +4.97
ARC-Challenge 25.17% 25.09% โˆ’0.08
WinoGrande 47.36% 50.20% +2.84
OpenBookQA 25.40% 27.20% +1.80
CommonsenseQA 20.31% 20.56% +0.25

Consistent improvements across reasoning benchmarks with a 2.7ร— parameter increase, demonstrating efficient scaling within the ultra-compact model regime.


We are happy to provide any additional evaluation details or run supplementary benchmarks if required. Thank you for maintaining this valuable resource for the SLM community.

Best regards,
Michelangelo Di Nicola
ThingAI Team

Axiomic Labs org

Hi, Thankyou for your submission! Im always happy to add more quality models to your
Ill verify and upload the benchmarks by end of (sydney) Day.

Axiomic Labs org
โ€ข
edited Jun 4

Its up! Impressive results, especially 135m on arithmark-2!
Out of interest what data was it trained on?

Thanks for adding them so quickly! Great to see the arithmark-2 results
The training data for both models:

open-web-math
smollm-corpus
proof-pile-2
the-stack-smol

Datdanboi25 changed discussion status to closed

Sign up or log in to comment