File size: 2,263 Bytes
4886f06
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e6997cf
4886f06
623f728
 
e6997cf
4886f06
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
e6997cf
 
4886f06
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
---
language:
- en
library_name: transformers
pipeline_tag: text-generation
datasets:
- epfml/FineWeb-HQ
- HuggingFaceFW/fineweb-edu
- mlfoundations/dclm-baseline-1.0
- HuggingFaceTB/smollm-corpus
- HuggingFaceTB/finemath
- AxiomicLabs/NPset-2-Python-Edu
tags:
- causal-lm
- base-model
- mixture-of-experts
- sequence-routing
- custom-code
- trust-remote-code
---

# SLMoE Test - 100%

This model has been released as https://huggingface.co/BananaMind/BananaMind-2-SLMoE

This is the **100% checkpoint** of an experimental sequence-routed base
language model. It is not instruction tuned.

## Architecture

| Field | Value |
|---|---:|
| Total parameters | 25,449,472 |
| Active parameters per sequence | 7,902,208 |
| Layers / hidden size | 8 / 256 |
| Query / KV heads | 8 / 2 |
| Head dimension | 32 |
| Expert paths | 64 total / 13 active |
| Expert intermediate size | 56 |
| Context | 4,096 |
| Vocabulary | 8,192, tied |

The router reads only the first 32 valid tokens and makes one global top-13
decision. Those expert identities are reused in every layer and stored in the
KV cache for the complete generated response. Routing does not change per
token. During pretraining, one packed 4,096-token sequence is the routing unit,
and loss on its routing prefix is masked to prevent future-token leakage.

## Training

| Field | Value |
|---|---:|
| Progress | 100% |
| Tokens seen | 59,999,518,720 |
| Target tokens | 60,000,000,000 |
| Hardware | 4 x NVIDIA H200 |
| Optimizer | Fused AdamW |
| Peak LR | 0.003 |
| Betas | 0.9, 0.95 |
| Weight decay | 0.1 |
| Precision | bfloat16 autocast |

| Token range | FineWeb-HQ | FineWeb-Edu | DCLM | Cosmopedia v2 | FineMath | NPSet-2 |
|---|---:|---:|---:|---:|---:|---:|
| 0.00B-12.00B | 35% | 30% | 20% | 10% | 4% | 1% |
| 12.00B-24.00B | 32% | 30% | 18% | 11% | 7% | 2% |
| 24.00B-39.00B | 28% | 29% | 16% | 13% | 11% | 3% |
| 39.00B-51.00B | 24% | 27% | 13% | 14% | 18% | 4% |
| 51.00B-60.00B | 20% | 24% | 9% | 15% | 26% | 6% |

## Usage

```python
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Banaxi-Tech/slmoe-test"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)
```