Blaze / README.md
spitfire4794's picture
Create README.md
4c203d0 verified
|
Raw
History Blame Contribute Delete
2.18 kB
---
language:
- en
tags:
- text-generation
- causal-lm
- custom-architecture
- slm
- small-language-model
license: mit
---
# Blaze (48.3M)
Blaze is a 48.3M parameter causal language model developed by SurjoLabs. It scores **15.45 on the Intelligence Index**, placing #1 in the sub-50M parameter category on the Open SLM Leaderboard.
The model uses XSA (orthogonal value-subtraction) attention with recurrent layer sharing, achieving an effective computational depth of 26 layers while storing only 14 physical layers.
---
## Architecture Specifications
| Parameter | Value |
| :--- | :--- |
| Total Parameters | 48,251,136 |
| Physical Layers | 14 (1 prelude + 12 recurrent + 1 coda) |
| Recurrent Passes | 2 (effective depth: 26 layers) |
| Hidden Size | 512 |
| Intermediate Size | 1536 |
| Attention Heads | 8 Query, 4 Key-Value (2:1 GQA) |
| Head Dimension | 64 |
| Vocabulary Size | 8,192 (tied embeddings) |
| Context Length | 1,024 tokens |
---
## Training & Checkpoint Selection
* **Total Tokens:** ~20.97B tokens (20,000 steps at 2^20 = 1,048,576 tokens/step)
* **Schedule:** WSD (Warmup-Stable-Decay) learning rate scheduler
* **Selected Checkpoint:** Checkpoint 19,500 achieved peak performance across benchmarks and is the official set of weights released in this repository.
---
## Benchmark Results
Evaluated 0-shot using normalized accuracy (acc_norm):
| Benchmark | Score |
| :--- | :--- |
| **PIQA** | 62.51% |
| **ARC-Easy** | 41.84% |
| **ArithMark-3.0** | 37.80% |
| **HellaSwag** | 31.84% |
| **ARC-Challenge** | 24.91% |
| **Intelligence Index** | **15.45** |
---
## Usage
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "SurjoLabs/Blaze"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
model_id,
trust_remote_code=True,
torch_dtype=torch.bfloat16,
).cuda()
prompt = "The speed of light is"
inputs = tokenizer(prompt, return_tensors="pt").to("cuda")
outputs = model.generate(**inputs, max_new_tokens=32)
print(tokenizer.decode(outputs[0], skip_special_tokens=True))
```
---
## License
MIT