File size: 8,411 Bytes
c2ccc8f 300239b c2ccc8f d8850c2 71b207e 45fa1c9 c2ccc8f a82591c 71b207e d8850c2 71b207e d8850c2 71b207e fd5fc66 a82591c c52af1b 300239b 0b15048 300239b 71b207e 4c9e96d 282b698 759c86f 282b698 759c86f f7cafa2 71b207e f7cafa2 4c9e96d 71b207e 56befbc 76b340d 56befbc 41b0cda 56befbc 41b0cda 56befbc 282b698 5434b4d 71b207e 45fa1c9 47c3a99 6d94972 47c3a99 6f9eb7e 9438344 6f9eb7e 56befbc 2b9bf3c 3447011 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 | ---
license: cc-by-4.0
language:
- en
pipeline_tag: text-generation
tags:
- text-generation-inference
- reasoning
- grpo
- qwen3
- atomight
- atomightv2-5
datasets:
- Logics-MLLM/Logics-STEM-SFT-Dataset-Open-1.6M
- MegaScience/MegaScience
- nvidia/OpenMathReasoning
- zwhe99/DeepMath-103K
- trl-lib/DeepMath-103K
- nvidia/OpenCodeInstruct
- nvidia/Nemotron-SFT-Competitive-Programming-v2
base_model:
- Qwen/Qwen3-1.7B
---
<div align="center">
<img src="NewAtomight.png" alt="Atomight v2 Logo" width="500" style="max-width: 100%;">
# Atomight V2.5 · 1.7B
**Reasoning-first · Zero benchmark contamination · Trained on a free Colab T4**




</div>
---
We are excited to announce and show you all, our most **powerful and capable model** in the current and *newest* Atomight family variant (*V2.5*).
> [!Note]
> Use this model when you want explicit chain-of-thought before the final answer — complex debugging, multi-step planning, agentic workflows, and math- or reasoning-heavy tasks.
> [!Note]
> ***Knowledge Span/Cutoff: Late 2024 / Early 2025 – Late 2025***
Other details for this model soon. Wait for further information and details.
---
## Quick Facts
| | |
|---|---|
| **Base model** | [Qwen/Qwen3-1.7B](https://huggingface.co/Qwen/Qwen3-1.7B) (Apache 2.0) |
| **Training method** | GRPO via LoRA/PEFT, merged to 16-bit |
| **Trained on** | Free-tier Google Colab T4 — no paid compute |
| **Training data** | ~2,000 curated samples per domain, 6 premium open datasets |
| **Contamination** | Zero — no benchmark data used in training |
| **License** | CC-BY-4.0 |
---
## Training Data
Curated (not scraped) from frontier 2025-era open datasets across STEM, science, math, and code:
| Domain | Dataset | License |
|---|---|---|
| STEM | [Logics-STEM-SFT-Dataset-Open-5.3M](https://huggingface.co/datasets/Logics-MLLM/Logics-STEM-SFT-Dataset-Open-5.3M) | Mixed (aggregated open sources) |
| Science | [MegaScience](https://huggingface.co/datasets/MegaScience/MegaScience) | CC-BY / Academic Use |
| Math | [OpenMathReasoning](https://huggingface.co/datasets/nvidia/OpenMathReasoning) | CC-BY-4.0 |
| Math | [DeepMath-103K](https://huggingface.co/datasets/zwhe99/DeepMath-103K) | CC-BY-4.0 |
| Code | [OpenCodeInstruct](https://huggingface.co/datasets/nvidia/OpenCodeInstruct) | CC-BY-4.0 |
| Code | [Nemotron-SFT-Competitive-Programming-v2](https://huggingface.co/datasets/nvidia/Nemotron-SFT-Competitive-Programming-v2) | NVIDIA Open Model License |
> [!Important]
> **No benchmark data was used in training.** None of the sources above overlap with MMLU, GSM8K, HumanEval, MBPP, HellaSwag, WinoGrande, ARC-Challenge, or TruthfulQA. Every score below was earned, not leaked.
---
## Benchmark results
These are the results of the benchmarks for *Atomight-V2.5-1.7B*, evaluated with `lm-evaluation-harness`.
| Benchmark | Metric | Score |
|---|---:|---:|
| MMLU | Accuracy | **55.68** |
| GSM8K | Accuracy | **69.60** |
| ARC-Challenge | Accuracy (normalized) | 43.00 |
| HellaSwag | Accuracy (normalized) | 60.43 |
| WinoGrande | Accuracy | 61.09 |
| TruthfulQA MC2 | Accuracy | 45.89 |
| HumanEval | Pass@1 | 40.24 |
| MBPP | Pass@1 | 42.80 |
<sub>All scores self-reported; no training data overlaps with these benchmarks — see Training Data section above.</sub>
---
## IIfSLM Benchmark Results
[IIfSLM](https://huggingface.co/datasets/NovatasticRoScript/IIfSLM-v1) (Intelligence Index for Small Language Models) is an open, contamination-resistant benchmark suite for the 0.5B–3B range — built for the whole small-model community to evaluate against, not exclusive to this model.
| Domain | Score | Details |
|---|---:|---|
| gsm8krefn | **86.96%** | 260/299 correct · greedy decoding · max_new_tokens=900 |
| arcchalrefn | **94.00%** | 235/250 correct · greedy decoding · max_new_tokens=1100 |
| humanevalrefn | In Progress | -- |
Methodology and full results are documented in the accompanying paper:
> NovatasticRoScript (2026). *IIfSLM: An Intelligence Index for Small Language Models.* Zenodo. [https://doi.org/10.5281/zenodo.21753925](https://doi.org/10.5281/zenodo.21753925)
See the [full IIfSLM dataset and methodology notes](https://huggingface.co/datasets/NovatasticRoScript/IIfSLM-v1) -- and feel free to run your own model against it too.
---
## How it compares with other small language models (we recommend verifying it, as the other data from other models came from a third-party sources)
Scores below for other models are drawn from their respective model cards / technical reports, not re-run by us. Provided for context only — evaluation harnesses and prompt formats differ across labs, so treat this as directional rather than exact.
| Model | Params | MMLU | GSM8K | HumanEval | MBPP | HellaSwag | WinoGrande |
|---|---:|---:|---:|---:|---:|---:|---:|
| **Atomight-V2.5-1.7B** | 1.7B | **55.68** | **69.60** | 40.24 | 42.80 | 60.43 | 61.09 |
| SmolLM2-1.7B (base) | 1.7B | 49.46 | 67.14 | 47.68 | 51.87 | 57.95 | 66.35 |
| Qwen2.5-1.5B (base) | 1.5B | 63.03 | 66.57 | 35.37 | 58.37 | 66.60 | 66.20 |
| Qwen2.5-1.5B-Instruct | 1.5B | 61.78 | 74.30 | 51.83 | 56.81 | — | — |
| InfiR-1B-Instruct | 1.0B | 50.22 | 70.90 | 58.54 | 56.03 | — | — |
| Llama-3.2-1B-Instruct | 1.0B | 46.27 | 47.90 | 39.63 | 49.03 | — | — |
| TinyLlama-1.1B | 1.1B | ~27 | ~9 | ~9 | ~27 | ~59 | ~61 |
**Where Atomight-V2.5-1.7B leads:** best-in-class GSM8K among comparable 1–2B base models, and MMLU well above SmolLM2-1.7B — achieved with a curated, contamination-free ~2K-per-category dataset rather than large-scale pretraining.
**Where it trails:** HumanEval and MBPP lag behind instruction-tuned peers like InfiR-1B-Instruct and Qwen2.5-1.5B-Instruct — expected, since Atomight-V2.5-1.7B is a base reasoning model without dedicated code SFT. HellaSwag also trails models trained on broader web/commonsense data.
---
## Quick Start
```python
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "NovatasticRoScript/Atomight-V2.5-1.7B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, device_map="auto")
messages = [
{"role": "system", "content": "You are a reasoning model. Think step-by-step inside <thinking> tags, then give your final answer inside <answer> tags."},
{"role": "user", "content": "If a train travels 60 miles in 45 minutes, what is its speed in mph?"}
]
inputs = tokenizer.apply_chat_template(messages, add_generation_prompt=True, return_tensors="pt").to(model.device)
outputs = model.generate(inputs, max_new_tokens=512)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
### Main citation
This model is a fine-tuned derivative of Qwen3-1.7B (Apache 2.0) and is released under CC-BY-4.0 in accordance with the licenses of its training datasets. If you use this model, kindly **credit `NovatasticRoScript/Atomight-V2.5-1.7B` and the base model/datasets listed above.** Part of the Atomight family — small models, curated data, no shortcuts.
### Other citacion/s
```
@misc{qwen3technicalreport,
title={Qwen3 Technical Report},
author={Qwen Team},
year={2025},
eprint={2505.09388},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2505.09388},
}
```
Training method — this model was trained using GRPO (Group Relative Policy Optimization), introduced in:
```
@article{deepseekmath2024,
title={DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models},
author={Shao, Zhihong and Wang, Peiyi and Zhu, Qihao and Xu, Runxin and Song, Junxiao and Zhang, Mingchuan and Li, Y.K. and Wu, Y. and Guo, Daya},
journal={arXiv preprint arXiv:2402.03300},
year={2024}
}
```
```
@misc{iifslm2026,
author = {NovatasticRoScript},
title = {IIfSLM: An Intelligence Index for Small Language Models — A Contamination-Resistant, Community-Driven Benchmark Suite for the 0.5B–3B Parameter Range},
year = {2026},
publisher = {Zenodo},
doi = {10.5281/zenodo.21753925},
url = {https://doi.org/10.5281/zenodo.21753925}
}
``` |