File size: 8,411 Bytes
c2ccc8f
300239b
c2ccc8f
 
 
d8850c2
 
71b207e
 
 
 
 
45fa1c9
 
 
 
 
 
 
 
 
 
c2ccc8f
a82591c
71b207e
d8850c2
71b207e
d8850c2
71b207e
 
 
 
 
 
 
 
 
 
 
 
 
fd5fc66
a82591c
c52af1b
 
 
300239b
0b15048
300239b
71b207e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
4c9e96d
282b698
759c86f
282b698
759c86f
f7cafa2
 
71b207e
 
f7cafa2
 
 
 
 
 
4c9e96d
71b207e
 
 
 
56befbc
 
76b340d
56befbc
 
 
 
41b0cda
 
56befbc
41b0cda
 
 
 
 
56befbc
 
 
282b698
5434b4d
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
71b207e
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
45fa1c9
 
47c3a99
 
6d94972
47c3a99
 
 
 
 
 
 
 
 
 
 
 
 
6f9eb7e
9438344
 
6f9eb7e
 
 
 
 
 
 
56befbc
2b9bf3c
 
 
 
 
 
 
 
 
3447011
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
---
license: cc-by-4.0
language:
- en
pipeline_tag: text-generation
tags:
- text-generation-inference
- reasoning
- grpo
- qwen3
- atomight
- atomightv2-5
datasets:
- Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M
- MegaScience/MegaScience
- nvidia/OpenMathReasoning
- zwhe99/DeepMath-103K
- trl-lib/DeepMath-103K
- nvidia/OpenCodeInstruct
- nvidia/Nemotron-SFT-Competitive-Programming-v2
base_model:
- Qwen/Qwen3-1.7B
---

<div align="center">

<img src="NewAtomight.png" alt="Atomight v2 Logo" width="500" style="max-width: 100%;">

# Atomight V2.5 · 1.7B

**Reasoning-first · Zero benchmark contamination · Trained on a free Colab T4**

![License](https://img.shields.io/badge/license-CC--BY--4.0-blue)
![Base](https://img.shields.io/badge/base-Qwen3--1.7B-orange)
![Method](https://img.shields.io/badge/method-GRPO-purple)
![Params](https://img.shields.io/badge/params-1.7B-green)

</div>

---

We are excited to announce and show you all, our most **powerful and capable model** in the current and *newest* Atomight family variant (*V2.5*).

> [!Note]
> Use this model when you want explicit chain-of-thought before the final answer — complex debugging, multi-step planning, agentic workflows, and math- or reasoning-heavy tasks.

> [!Note]
> ***Knowledge Span/Cutoff: Late 2024 / Early 2025 – Late 2025***

Other details for this model soon. Wait for further information and details.

---

## Quick Facts

| | |
|---|---|
| **Base model** | [Qwen/Qwen3-1.7B](https://huggingface.co/Qwen/Qwen3-1.7B) (Apache 2.0) |
| **Training method** | GRPO via LoRA/PEFT, merged to 16-bit |
| **Trained on** | Free-tier Google Colab T4 — no paid compute |
| **Training data** | ~2,000 curated samples per domain, 6 premium open datasets |
| **Contamination** | Zero — no benchmark data used in training |
| **License** | CC-BY-4.0 |

---

## Training Data

Curated (not scraped) from frontier 2025-era open datasets across STEM, science, math, and code:

| Domain | Dataset | License |
|---|---|---|
| STEM | [Logics-STEM-SFT-Dataset-Open-5.3M](https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-5.3M) | Mixed (aggregated open sources) |
| Science | [MegaScience](https://huggingface.co/datasets/MegaScience/MegaScience) | CC-BY / Academic Use |
| Math | [OpenMathReasoning](https://huggingface.co/datasets/nvidia/OpenMathReasoning) | CC-BY-4.0 |
| Math | [DeepMath-103K](https://huggingface.co/datasets/zwhe99/DeepMath-103K) | CC-BY-4.0 |
| Code | [OpenCodeInstruct](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) | CC-BY-4.0 |
| Code | [Nemotron-SFT-Competitive-Programming-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Competitive-Programming-v2) | NVIDIA Open Model License |

> [!Important]
> **No benchmark data was used in training.** None of the sources above overlap with MMLU, GSM8K, HumanEval, MBPP, HellaSwag, WinoGrande, ARC-Challenge, or TruthfulQA. Every score below was earned, not leaked.

---

## Benchmark results

These are the results of the benchmarks for *Atomight-V2.5-1.7B*, evaluated with `lm-evaluation-harness`.

| Benchmark | Metric | Score |
|---|---:|---:|
| MMLU | Accuracy | **55.68** |
| GSM8K | Accuracy | **69.60** |
| ARC-Challenge | Accuracy (normalized) | 43.00 |
| HellaSwag | Accuracy (normalized) | 60.43 |
| WinoGrande | Accuracy | 61.09 |
| TruthfulQA MC2 | Accuracy | 45.89 |
| HumanEval | Pass@1 | 40.24 |
| MBPP | Pass@1 | 42.80 |

<sub>All scores self-reported; no training data overlaps with these benchmarks — see Training Data section above.</sub>

---

## IIfSLM Benchmark Results

[IIfSLM](https://huggingface.co/datasets/NovatasticRoScript/IIfSLM-v1) (Intelligence Index for Small Language Models) is an open, contamination-resistant benchmark suite for the 0.5B–3B range — built for the whole small-model community to evaluate against, not exclusive to this model.

| Domain | Score | Details |
|---|---:|---|
| gsm8krefn | **86.96%** | 260/299 correct · greedy decoding · max_new_tokens=900 |
| arcchalrefn | **94.00%** | 235/250 correct · greedy decoding · max_new_tokens=1100 |
| humanevalrefn | In Progress | -- |

Methodology and full results are documented in the accompanying paper:

> NovatasticRoScript (2026). *IIfSLM: An Intelligence Index for Small Language Models.* Zenodo. [https://doi.org/10.5281/zenodo.21753925](https://doi.org/10.5281/zenodo.21753925)

See the [full IIfSLM dataset and methodology notes](https://huggingface.co/datasets/NovatasticRoScript/IIfSLM-v1) -- and feel free to run your own model against it too. 

---

## How it compares with other small language models (we recommend verifying it, as the other data from other models came from a third-party sources) 

Scores below for other models are drawn from their respective model cards / technical reports, not re-run by us. Provided for context only — evaluation harnesses and prompt formats differ across labs, so treat this as directional rather than exact.

| Model | Params | MMLU | GSM8K | HumanEval | MBPP | HellaSwag | WinoGrande |
|---|---:|---:|---:|---:|---:|---:|---:|
| **Atomight-V2.5-1.7B** | 1.7B | **55.68** | **69.60** | 40.24 | 42.80 | 60.43 | 61.09 |
| SmolLM2-1.7B (base) | 1.7B | 49.46 | 67.14 | 47.68 | 51.87 | 57.95 | 66.35 |
| Qwen2.5-1.5B (base) | 1.5B | 63.03 | 66.57 | 35.37 | 58.37 | 66.60 | 66.20 |
| Qwen2.5-1.5B-Instruct | 1.5B | 61.78 | 74.30 | 51.83 | 56.81 | — | — |
| InfiR-1B-Instruct | 1.0B | 50.22 | 70.90 | 58.54 | 56.03 | — | — |
| Llama-3.2-1B-Instruct | 1.0B | 46.27 | 47.90 | 39.63 | 49.03 | — | — |
| TinyLlama-1.1B | 1.1B | ~27 | ~9 | ~9 | ~27 | ~59 | ~61 |

**Where Atomight-V2.5-1.7B leads:** best-in-class GSM8K among comparable 1–2B base models, and MMLU well above SmolLM2-1.7B — achieved with a curated, contamination-free ~2K-per-category dataset rather than large-scale pretraining.

**Where it trails:** HumanEval and MBPP lag behind instruction-tuned peers like InfiR-1B-Instruct and Qwen2.5-1.5B-Instruct — expected, since Atomight-V2.5-1.7B is a base reasoning model without dedicated code SFT. HellaSwag also trails models trained on broader web/commonsense data.

---

## Quick Start

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "NovatasticRoScript/Atomight-V2.5-1.7B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")

messages = [
    {"role": "system", "content": "You are a reasoning model. Think step-by-step inside <thinking> tags, then give your final answer inside <answer> tags."},
    {"role": "user", "content": "If a train travels 60 miles in 45 minutes, what is its speed in mph?"}
]

inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
### Main citation

This model is a fine-tuned derivative of Qwen3-1.7B (Apache 2.0) and is released under CC-BY-4.0 in accordance with the licenses of its training datasets. If you use this model, kindly **credit `NovatasticRoScript/Atomight-V2.5-1.7B` and the base model/datasets listed above.** Part of the Atomight family — small models, curated data, no shortcuts.

### Other citacion/s

```
@misc{qwen3technicalreport,
      title={Qwen3 Technical Report}, 
      author={Qwen Team},
      year={2025},
      eprint={2505.09388},
      archivePrefix={arXiv},
      primaryClass={cs.CL},
      url={https://arxiv.org/abs/2505.09388}, 
}
```
Training method — this model was trained using GRPO (Group Relative Policy Optimization), introduced in:

```
@article{deepseekmath2024,
  title={DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models},
  author={Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Zhang, Mingchuan and Li, Y.K. and Wu, Y. and Guo, Daya},
  journal={arXiv preprint arXiv:2402.03300},
  year={2024}
}
```
```
@misc{iifslm2026,
  author       = {NovatasticRoScript},
  title        = {IIfSLM: An Intelligence Index for Small Language Models — A Contamination-Resistant, Community-Driven Benchmark Suite for the 0.5B–3B Parameter Range},
  year         = {2026},
  publisher    = {Zenodo},
  doi          = {10.5281/zenodo.21753925},
  url          = {https://doi.org/10.5281/zenodo.21753925}
}
```