SLMoE Test - 100%

This model has been released as https://huggingface.co/BananaMind/BananaMind-2-SLMoE

This is the 100% checkpoint of an experimental sequence-routed base language model. It is not instruction tuned.

Architecture

Field Value
Total parameters 25,449,472
Active parameters per sequence 7,902,208
Layers / hidden size 8 / 256
Query / KV heads 8 / 2
Head dimension 32
Expert paths 64 total / 13 active
Expert intermediate size 56
Context 4,096
Vocabulary 8,192, tied

The router reads only the first 32 valid tokens and makes one global top-13 decision. Those expert identities are reused in every layer and stored in the KV cache for the complete generated response. Routing does not change per token. During pretraining, one packed 4,096-token sequence is the routing unit, and loss on its routing prefix is masked to prevent future-token leakage.

Training

Field Value
Progress 100%
Tokens seen 59,999,518,720
Target tokens 60,000,000,000
Hardware 4 x NVIDIA H200
Optimizer Fused AdamW
Peak LR 0.003
Betas 0.9, 0.95
Weight decay 0.1
Precision bfloat16 autocast
Token range FineWeb-HQ FineWeb-Edu DCLM Cosmopedia v2 FineMath NPSet-2
0.00B-12.00B 35% 30% 20% 10% 4% 1%
12.00B-24.00B 32% 30% 18% 11% 7% 2%
24.00B-39.00B 28% 29% 16% 13% 11% 3%
39.00B-51.00B 24% 27% 13% 14% 18% 4%
51.00B-60.00B 20% 24% 9% 15% 26% 6%

Usage

from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "Banaxi-Tech/slmoe-test"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(model_id, trust_remote_code=True)
Downloads last month
-
Safetensors
Model size
27.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train Banaxi-Tech/slmoe-test