Text Generation
Transformers
Safetensors
PyTorch
cma
custom_code
causal-lm
small-language-model
generalist
4k-tokenizer
Instructions to use User01110/cma-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use User01110/cma-mini with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="User01110/cma-mini", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("User01110/cma-mini", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use User01110/cma-mini with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "User01110/cma-mini" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/cma-mini", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/User01110/cma-mini
- SGLang
How to use User01110/cma-mini with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "User01110/cma-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/cma-mini", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "User01110/cma-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/cma-mini", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use User01110/cma-mini with Docker Model Runner:
docker model run hf.co/User01110/cma-mini
File size: 7,603 Bytes
b184320 21efcb1 b184320 21efcb1 b184320 21efcb1 b184320 21efcb1 b184320 21efcb1 b184320 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 148 149 150 151 152 153 154 155 156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 | ---
library_name: transformers
pipeline_tag: text-generation
tags:
- custom_code
- causal-lm
- text-generation
- pytorch
- cma
- small-language-model
- generalist
- 4k-tokenizer
datasets:
- HuggingFaceFW/fineweb_edu_100BT-shuffled
- HuggingFaceTB/smollm-corpus
- epfml/FineWeb-HQ
- HuggingFaceTB/finemath
---
# CMA Mini
Evaluated training checkpoint from a 5.09M-parameter pre-CMA
generalist language model using the third-party
AxiomicLabs GPT-S 4,096-token tokenizer. It has no place embeddings, role
embeddings, or inference-time equation detection. It was recorded at step
60,000 with WikiText normalized BPB 1.4127.
## Loading
This is a custom Transformers architecture. `trust_remote_code=True` is
required because stock Hugging Face model classes do not implement CMA or this
model's exact rotary convention.
```bash
pip install "torch>=2.5" "transformers>=4.50" safetensors
```
```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "User01110/cma-mini"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo,
trust_remote_code=True,
torch_dtype=dtype,
).to(device).eval()
inputs = tokenizer("The process of photosynthesis", return_tensors="pt").to(device)
output = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(output[0], skip_special_tokens=True))
```
Checkpoint tensors are stored in bfloat16. Pass `torch_dtype=torch.float32`
when an FP32 runtime is required; every stored BF16 value widens exactly to
FP32, though the pre-export FP32 master-weight mantissa cannot be reconstructed.
The model intentionally sets `use_cache=False`: generation is standard
Transformers generation, but the visible context is recomputed for every new
token because this experimental architecture does not implement a KV cache.
## Architecture
- Parameters: 5,094,112, with tied input/output embeddings
- Weights: native bfloat16 safetensors (`model.safetensors`); no `.bin` weights
- Runtime: native SDPA for token attention and dense native PyTorch CMA routing
- Tokenizer: `AxiomicLabs/GPT-S-5M` at revision
`275b9c3ca78736bf6aeb154c7e2d5f5764fe9035`
- Vocabulary: 4,096 third-party tokens
- Parameter allocation: 884,736 tied embedding parameters and
4,209,376 non-embedding parameters
- Supported/exported context: 1,024 tokens
- Training block length: 1,024 tokens
- Standalone prompt tokenization automatically prepends the native BOS token
- Width/layers: 216 / 10
- Token-attention heads: 6 query, 2 KV
- CMA reshapes each token into 9 channel chunks of
24 channels, with 3 routing heads and expansion
2
- Every routing head computes an exact dense
9x9 softmax interaction
graph; no attention edge is removed, factorized, or approximated
- Queries combine the local chunk, a learned chunk identity, and a global
token summary; keys are conditioned on the expanded value representation
- Routing is a signed residual update
`V + tanh(g) * (Attention(Q,K,V) - V)`, so CMA can refine, reverse, or retain
each chunk instead of overwriting it
- The learned diagonal bias starts at 90%
identity mass, providing a stable path while learned cross-chunk routing grows
- The expanded routed values are modulated by a SiLU content gate before the
zero-initialized output projection
- The output projection starts at zero for stable residual initialization
- Contiguous-half RoPE without scaling
- No task-specific model features or inference-time benchmark handling
## Intended use and limitations
This is a small base language model released for architecture research,
representation analysis, and controlled comparisons. It is not instruction
tuned and should not be treated as a factual authority or used for consequential
decisions. Its 5.09M-parameter scale, English-weighted web mixture,
1,024-token context, and cache-free generation materially limit
capability and throughput. The usual web-corpus biases and inaccuracies remain.
## Tokenizer provenance
The tokenizer and its 4,096-token vocabulary were **not
created or owned by the CMA model author**. They are reused from the public
[`AxiomicLabs/GPT-S-5M`](https://huggingface.co/AxiomicLabs/GPT-S-5M)
repository at the exact revision listed above, whose repository metadata
identifies Axiomic Labs as the publisher and Apache-2.0 as the license. No
claim of tokenizer ownership beyond that public attribution is made here. The
exported copy preserves its vocabulary and tokenization pipeline; CMA only
configures its existing `<bos>` token to be prepended automatically and raises
the exported model/tokenizer context limit to 1,024 tokens.
## Optimization
- Training budget: 31,457,280,000 tokens over 60,000 updates
- Effective batch: 524,288 tokens per update
- Native microbatch: 256 sequences x 1,024 tokens,
accumulated 2 times
- Runtime implementation: explicit BF16 residual/CMA activations, dense
9x9 channel-routing matmuls, and full-model `torch.compile` enabled by default
- Runtime: 2 resumable 30,000-update sessions
with exact full-precision optimizer/RNG/data-stream recovery checkpoints
- Learning rate: 750-update linear warmup, flat at
2.5e-03 through update 30,000, linear decline to
1.5e-03 at update 50,000, then cosine toward
zero through update 60,000; original-scaling Muon follows the same
multiplier from its 3.0e-02 peak
- The final configured update is positive; update 60,001 is exactly zero
- Official PyTorch Muon with original matrix-shape scaling for hidden matrices;
AdamW for embeddings and remaining parameters
## Training mixture
- FineWeb-Edu 100BT shuffled: 55%
- Cosmopedia v2: 25%
- FineWeb-HQ: 10%
- FineMath 4+: 10%
FineWeb-Edu supplies the primary educational web text, Cosmopedia supplies
synthetic textbook-style coverage, FineWeb-HQ contributes model-filtered,
knowledge-rich general web text, and FineMath-4+ supplies mathematical
reasoning as ordinary causal-language-model text. The mixture remains fixed
for the complete run.
There are no benchmark labels or benchmark-specific preprocessing. Dataset
revisions are pinned and the same fixed mixture is maintained across both
runtime sessions.
## Zero-shot evaluation at step 60,000
The four lm-eval tasks use normalized accuracy when supplied by lm-eval
0.4.12, with native bfloat16 weights and float32
likelihood softmax. ArithMark-3 uses its official primary `acc_norm` metric:
mean continuation-token log likelihood, with context and continuation tokenized
separately after one native BOS prefix. Autocast is not used for evaluation. The four
lm-eval benchmark contexts use the tokenizer's native BOS behavior.
**Open SLM Intelligence Index: 6.84**
**Open SLM Average: 35.26%**
| Benchmark | Accuracy |
|---|---:|
| HellaSwag | 27.86% |
| ARC-Easy | 33.50% |
| ARC-Challenge | 22.70% |
| PIQA | 57.29% |
| ArithMark-3 (`acc_norm`) | 27.80% |
The Average equally weights HellaSwag, combined ARC (the mean of ARC-Easy and
ARC-Challenge), PIQA, and ArithMark-3. The Intelligence Index first maps random
chance to 0 and perfect accuracy to 100, then weights HellaSwag, combined ARC,
and PIQA at 1.0 and ArithMark-3 at 0.65, matching Open SLM Leaderboard revision
`2fbaaa164009c6c0d201ddf245878f5942cc8628`.
WikiText-103 validation at this step: loss 3.1243, perplexity
22.74, normalized BPB 1.4127 over 359,037
scored tokens and 1,145,591 normalized UTF-8 bytes, using one initial BOS,
1,024-token windows, and a 512-token stride.
|