Text Generation
Transformers
Safetensors
English
qwen3
small
tiny
supra
supra2
efficient
text-generation-inference
Instructions to use SupraLabs/Supra2-Medium-Base with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use SupraLabs/Supra2-Medium-Base with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="SupraLabs/Supra2-Medium-Base")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("SupraLabs/Supra2-Medium-Base") model = AutoModelForCausalLM.from_pretrained("SupraLabs/Supra2-Medium-Base", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use SupraLabs/Supra2-Medium-Base with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "SupraLabs/Supra2-Medium-Base" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SupraLabs/Supra2-Medium-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/SupraLabs/Supra2-Medium-Base
- SGLang
How to use SupraLabs/Supra2-Medium-Base with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "SupraLabs/Supra2-Medium-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SupraLabs/Supra2-Medium-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "SupraLabs/Supra2-Medium-Base" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "SupraLabs/Supra2-Medium-Base", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use SupraLabs/Supra2-Medium-Base with Docker Model Runner:
docker model run hf.co/SupraLabs/Supra2-Medium-Base
File size: 8,117 Bytes
14b9d0a 942b8f3 14b9d0a 942b8f3 d6d70f8 942b8f3 dc2d1ea a7534a1 942b8f3 d6d70f8 63560e2 942b8f3 90f7f43 942b8f3 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 | ---
license: apache-2.0
viewer: false
datasets:
- HuggingFaceFW/fineweb-edu
language:
- en
pipeline_tag: text-generation
library_name: transformers
tags:
- small
- tiny
- supra
- supra2
- qwen3
- efficient
---
<h1 align="center">Supra2-Medium Base</h1>
<p align="center">
Ultra-efficient base model • 25M Parameters • 1K Context
</p>

**Supra2-Medium Base** is a 25M-parameter decoder-only language model pretrained from scratch by **SupraLabs** on 20B tokens of English web text. It uses the **Qwen3** architecture with a custom 16,384-token tokenizer.
This is a **base model**. It has *not* been instruction-tuned, chat-tuned, or aligned in any way.
At only 25 million parameters, Supra2-Medium demonstrates that meaningful language modeling can be achieved with extreme parameter efficiency—trained at ~800 tokens per parameter, which is significantly higher than typical pretraining ratios. This makes it ideal for research into data-efficient scaling and ultra-lightweight deployments.
---
## **Let the model speak - a sample**
Prompt: "Artificial intelligence (AI) is "
Completion:
```plaintext
Artificial intelligence (AI) is espoused by the AI community.
The AI community is a group of people who are interested in AI and are interested in the use of AI in the field of AI.
The goal of AI is to improve the quality of life of people in the field.
The aim of AI is the development of AI and the application of AI in a society.
The purpose of AI is that it can be used to improve the performance of the society.
It is a technology that is used to improve human intelligence.
The technology is used to make the human intelligence.
```
## **Evaluation & Benchmarks**
All benchmarks were evaluated using the EleutherAI LM-Eval Harness.
| Model | PIQA (acc\_norm) | HellaSwag (acc\_norm) | ARC-Easy (acc\_norm) | ARC-Challenge (acc\_norm) |
| ----- | :---: | :---: | :---: | :---: |
| Supra-50M-Base (50M) | 0.62 | 0.32 | 0.46 | 0.25 |
| **Supra2-Medium-Base (25M)** | **59.14** | **29.29** | **41.84** | **23.72** |
| Supra2-100M-Base (100M) | 0.65 | 0.36 | 0.48 | 0.25 |
**Final Train Loss: 3.2469** (no eval loss available)
*Note: The model shows strong performance relative to its size, particularly given the high token-per-parameter ratio. It even strongly competes with our previous 50M model!*
---
## **Model Details**
| | |
| ----- | ----- |
| **Developed by** | SupraLabs |
| **Model type** | Causal decoder-only transformer (Qwen3) |
| **Language** | English |
| **Parameters** | 25.37M total / \~20M non-embedding |
| **Training tokens** | 20B (~800 tokens per parameter) |
| **Context length** | 1,024 |
| **Precision** | bfloat16 |
| **License** | Apache 2.0 |
### **Architecture**
| Hyperparameter | Value |
| ----- | ----- |
| Hidden size | 512 |
| Layers | 7 |
| Attention heads | 8 (MHA) |
| Head dim | 64 |
| Intermediate size (SwiGLU) | 896 |
| Vocab size | 16,384 |
| Positional encoding | RoPE θ=10,000 |
| Normalization | RMSNorm, ε=10⁻⁶ |
| Tied embeddings | Yes |
| Attention implementation | SDPA |
| Architecture ID | `qwen3-d07-h512-i0896` |
---
## **Training Data**
| Source | Share | Approx. tokens |
| ----- | ----- | ----- |
| `HuggingFaceFW/fineweb-edu` (`sample-100BT`) | 100% | 20B |
Documents were tokenized with the custom `tokenizer_16k`, concatenated into a flat `uint16` token stream, and packed into contiguous 1,024-token chunks (no padding, no document masking — sequences may cross document boundaries).
---
## **Training Procedure**
| Setting | Value |
| ----- | ----- |
| Optimizer | AdamW (fused), β₁=0.9, β₂=0.95, ε=10⁻⁸ |
| Peak learning rate | 3×10⁻³ |
| LR schedule | Cosine with min LR (10% of peak) |
| Warmup steps | 1,000 |
| Micro batch size | 32 |
| Gradient accumulation | 4 |
| Effective batch | 256 sequences \= **262,144 tokens/step** |
| Weight decay | 0.1 |
| Gradient clipping | 1.0 |
| Mixed precision | BF16 \+ TF32 |
| Hardware | 2× GPU (DDP): RTX 5060 Ti 16GB + RTX 5060 8GB |
| Training optimizations | Liger Kernel, DDP bucketing (10MB), gradient\_as\_bucket\_view |
### **Key Design Decisions**
* **High token-per-parameter ratio**: At 800 tokens/param, this model pushes the boundaries of data efficiency for small models
* **Compact vocabulary**: 16K vocab reduces embedding overhead while maintaining coverage
* **Tied embeddings**: Reduces parameter count and improves training stability
* **Multi-head attention (MHA)**: Unlike larger Supra2 models using GQA, this architecture uses standard MHA for simplicity at this scale
* **No sliding window**: Full attention within the 1K context
---
## **Usage**
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "SupraLabs/Supra2-Medium"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
torch_dtype=torch.bfloat16,
device_map="auto",
)
model.eval()
prompt = "The future of artificial intelligence"
inputs = tokenizer(prompt, return_tensors="pt").to(model.device)
with torch.no_grad():
out = model.generate(
**inputs,
max_new_tokens=512,
do_sample=True,
temperature=0.2,
top_p=0.85,
top_k=25,
no_repeat_ngram_size=3,
repetition_penalty=1.2,
)
print(tokenizer.decode(out[0], skip_special_tokens=True))
```
### **Tokenizer notes**
The model uses a custom 16,384-token vocabulary optimized for English web text. The tokenizer is shared across the Supra2 family's smaller models for consistency and efficient multi-task fine-tuning.
---
## **Intended Use**
**Intended:**
* Research on extreme parameter efficiency and data-efficient pretraining
* Ultra-lightweight edge deployments where memory is severely constrained
* Educational use for understanding transformer architectures at minimal scale
* Starting point for domain-specific fine-tuning when compute resources are limited
* Ablation studies on small model behavior and scaling laws
**Not intended:**
* Production deployment for critical applications
* Factual question answering or knowledge-intensive tasks
* Long-form coherent generation beyond a few paragraphs
* Non-English text (trained exclusively on English)
* Any application requiring safety guarantees or alignment
---
## **Limitations and Bias**
* **Very small.** At 25M parameters, the model has extremely limited capacity. Expect frequent hallucinations, factual errors, repetitive outputs, and poor reasoning.
* **Base model.** No RLHF, no safety tuning, no refusal behavior. It will continue whatever text you give it, including harmful or offensive prompts.
* **Web-derived data.** FineWeb-Edu is a filtered CommonCrawl derivative and carries the biases, stereotypes, and factual errors of the open web.
* **Short context.** Trained exclusively at 1,024 tokens. Extrapolation beyond this length is untested and likely degraded.
* **No document masking.** Attention could cross document boundaries within a packed chunk, which slightly blurs document independence.
* **English only.** Performance on other languages will be poor to non-existent.
* **High token-per-parameter ratio.** While efficient, training at 800 tok/param means the model may be under-trained compared to models trained at lower ratios (e.g., 300 tok/param).
---
## **Performance Characteristics**
Despite its tiny size, Supra2-Medium achieves surprisingly coherent short-form generation. Key observations:
* **Strengths**: Basic grammar, simple factoids, short completions, pattern matching
* **Weaknesses**: Complex reasoning, arithmetic, long-range coherence, factual accuracy, nuanced understanding
* **Best use case**: Generating 1-3 sentence completions, simple text transformations, educational demonstrations
---
Future work includes instruction-tuned variants and exploration of even more efficient architectures at this scale.
*© SupraLabs 2026* |