Sol Nano
Sol Nano is a 2,895,188-parameter language model trained on 5 billion tokens. It uses the Sol Lite SolForCausalLM architecture with TN-Gram memory and a 1,024-entry tokenizer.
Model
| Setting | Value |
|---|---|
| Parameters | 2,895,188 |
| TN-Gram parameters | 209,748 |
| Hidden width | 128 |
| Context | 512 tokens |
| Vocabulary | 1,024 |
| Stored blocks / effective applications | 10 / 14 |
| Query heads / KV heads | 4 / 2 |
| FFN width | 536 |
| Training tokens | 5,000,000,000 |
| Optimizer updates | 19,074 |
| Weights | FP32 safetensors |
The blocks use causal grouped-query attention, RoPE, QK normalization, and loop conditioning. Token embeddings share weights with the output head. TN-Gram provides factorized local memory for orders 2-5.
Use
Load the included PyTorch implementation in a CUDA environment with Triton and FlexAttention support. Install huggingface_hub, tokenizers, and safetensors alongside PyTorch.
import os
import sys
from pathlib import Path
import torch
from huggingface_hub import snapshot_download
from safetensors.torch import load_file
from tokenizers import Tokenizer
os.environ["SOL_NANO_ATTENTION"] = "triton"
model_dir = Path(snapshot_download("solintellegence/sol-nano"))
sys.path.insert(0, str(model_dir))
from modeling_sol_lite import SolForCausalLM, variant_config
model = SolForCausalLM(variant_config("sol_nano_2p9m_tn_gram"))
model.load_state_dict(load_file(str(model_dir / "model.safetensors")), strict=True)
model = model.cuda().eval()
tokenizer = Tokenizer.from_file(str(model_dir / "tokenizer.json"))
prompt = "The sum of 12 and 7 is"
ids = tokenizer.encode(prompt, add_special_tokens=False).ids
inputs = torch.tensor([ids], dtype=torch.long, device="cuda")
with torch.inference_mode():
logits = model(inputs) # [batch, sequence, vocabulary]
next_id = logits[0, -1].argmax().item()
print(tokenizer.decode([next_id]))
Training
Fused AdamW trained the model in BF16 with FP32 optimizer states on one RTX PRO 6000 Blackwell Server Edition. Each full optimizer update contained 512 sequences of 512 tokens. CPU workers streamed and tokenized the sources while the GPU trained; each update included the complete scheduled mixture.
The peak learning rate was 0.001. WSD used a linear warmup over the first 2% of optimizer steps, a stable rate through 90%, and a linear decay to zero over the final 10%.
| Phase | FineWeb-Edu | FineMath | OpenWebMath | Generated math | Procedural | Physical science | Code |
|---|---|---|---|---|---|---|---|
| Opening, about 0-1.333B tokens | 65% | 7.5% | 4.5% | 3% | 12% | 4% | 4% |
| Main, about 1.4-4.5B tokens | 45% | 20% | 12% | 8% | 8% | 3% | 4% |
| Final 10% of optimizer steps | 30% | 30% | 20% | 10% | 4% | 2% | 4% |
A 66.85M-token ramp connects the opening and main phases. The final phase uses FineWeb-Edu scores of at least 3.5 and FineMath scores of at least 4.5. Procedural text is filtered from Cosmopedia-v2, physical-science text from FineWeb-Edu, and code from CoRNStack Python. Exact boundaries and source settings are in run.json.
Training used PyTorch 2.11.0+cu130 and Triton 3.6.0.
Evaluation
| Benchmark | Examples | Normalized accuracy |
|---|---|---|
| HellaSwag | 10,042 | 28.40% |
| ARC-Easy | 2,376 | 32.07% |
| ARC-Challenge | 1,172 | 21.16% |
| PIQA | 1,838 | 53.92% |
| ArithMark-3 | 1,000 | 33.80% |
The Axiomic Labs Open SLM Intelligence Index is 6.0684. These are full zero-shot results using LM Evaluation Harness 0.4.12 and the official ArithMark-3.0 dataset. Scoring used float32 and a 512-token context with PyTorch 2.14.0+cu130. No candidate request required truncation.
The index follows the published methodology. The results have not been independently verified by Axiomic Labs. Full-precision scores and checkpoint hashes are in evaluation/summary.json.
Files and limits
The repository contains model.safetensors, the matching tokenizer, modeling_sol_lite.py, configuration, and training metadata. The safetensors file is 11,591,920 bytes. Optimizer state is not included.
Nano is a base model for text completion. It has not been instruction-tuned, and its answers can be incorrect. The benchmark scores measure multiple-choice likelihood accuracy; they do not establish reliable free-form problem solving.
- Downloads last month
- 11
