Instructions to use User01110/cma-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use User01110/cma-mini with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="User01110/cma-mini", trust_remote_code=True)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("User01110/cma-mini", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use User01110/cma-mini with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "User01110/cma-mini" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/cma-mini", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker
docker model run hf.co/User01110/cma-mini
- SGLang
How to use User01110/cma-mini with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "User01110/cma-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/cma-mini", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "User01110/cma-mini" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/cma-mini", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }' - Docker Model Runner
How to use User01110/cma-mini with Docker Model Runner:
docker model run hf.co/User01110/cma-mini
CMA Mini
Evaluated training checkpoint from a 5.09M-parameter pre-CMA generalist language model using the third-party AxiomicLabs GPT-S 4,096-token tokenizer. It has no place embeddings, role embeddings, or inference-time equation detection. It was recorded at step 50,000 with WikiText normalized BPB 1.4397.
Loading
This is a custom Transformers architecture. trust_remote_code=True is
required because stock Hugging Face model classes do not implement CMA or this
model's exact rotary convention.
pip install "torch>=2.5" "transformers>=4.50" safetensors
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
repo = "User01110/cma-mini"
device = "cuda" if torch.cuda.is_available() else "cpu"
dtype = torch.bfloat16 if device == "cuda" else torch.float32
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
repo,
trust_remote_code=True,
torch_dtype=dtype,
).to(device).eval()
inputs = tokenizer("The process of photosynthesis", return_tensors="pt").to(device)
output = model.generate(**inputs, max_new_tokens=64)
print(tokenizer.decode(output[0], skip_special_tokens=True))
Checkpoint tensors are stored in bfloat16. Pass torch_dtype=torch.float32
when an FP32 runtime is required; every stored BF16 value widens exactly to
FP32, though the pre-export FP32 master-weight mantissa cannot be reconstructed.
The model intentionally sets use_cache=False: generation is standard
Transformers generation, but the visible context is recomputed for every new
token because this experimental architecture does not implement a KV cache.
Architecture
- Parameters: 5,094,112, with tied input/output embeddings
- Weights: native bfloat16 safetensors (
model.safetensors); no.binweights - Runtime: native SDPA for token attention and dense native PyTorch CMA routing
- Tokenizer:
AxiomicLabs/GPT-S-5Mat revision275b9c3ca78736bf6aeb154c7e2d5f5764fe9035 - Vocabulary: 4,096 third-party tokens
- Parameter allocation: 884,736 tied embedding parameters and 4,209,376 non-embedding parameters
- Supported/exported context: 1,024 tokens
- Training block length: 1,024 tokens
- Standalone prompt tokenization automatically prepends the native BOS token
- Width/layers: 216 / 10
- Token-attention heads: 6 query, 2 KV
- CMA reshapes each token into 9 channel chunks of 24 channels, with 3 routing heads and expansion 2
- Every routing head computes an exact dense 9x9 softmax interaction graph; no attention edge is removed, factorized, or approximated
- Queries combine the local chunk, a learned chunk identity, and a global token summary; keys are conditioned on the expanded value representation
- Routing is a signed residual update
V + tanh(g) * (Attention(Q,K,V) - V), so CMA can refine, reverse, or retain each chunk instead of overwriting it - The learned diagonal bias starts at 90% identity mass, providing a stable path while learned cross-chunk routing grows
- The expanded routed values are modulated by a SiLU content gate before the zero-initialized output projection
- The output projection starts at zero for stable residual initialization
- Contiguous-half RoPE without scaling
- No task-specific model features or inference-time benchmark handling
Intended use and limitations
This is a small base language model released for architecture research, representation analysis, and controlled comparisons. It is not instruction tuned and should not be treated as a factual authority or used for consequential decisions. Its 5.09M-parameter scale, English-weighted web mixture, 1,024-token context, and cache-free generation materially limit capability and throughput. The usual web-corpus biases and inaccuracies remain.
Tokenizer provenance
The tokenizer and its 4,096-token vocabulary were not
created or owned by the CMA model author. They are reused from the public
AxiomicLabs/GPT-S-5M
repository at the exact revision listed above, whose repository metadata
identifies Axiomic Labs as the publisher and Apache-2.0 as the license. No
claim of tokenizer ownership beyond that public attribution is made here. The
exported copy preserves its vocabulary and tokenization pipeline; CMA only
configures its existing <bos> token to be prepended automatically and raises
the exported model/tokenizer context limit to 1,024 tokens.
Optimization
- Training budget: 31,457,280,000 tokens over 60,000 updates
- Effective batch: 524,288 tokens per update
- Native microbatch: 256 sequences x 1,024 tokens, accumulated 2 times
- Runtime implementation: explicit BF16 residual/CMA activations, dense
9x9 channel-routing matmuls, and full-model
torch.compileenabled by default - Runtime: 2 resumable 30,000-update sessions with exact full-precision optimizer/RNG/data-stream recovery checkpoints
- Learning rate: 750-update linear warmup, flat at 2.5e-03 through update 30,000, linear decline to 1.5e-03 at update 50,000, then cosine toward zero through update 60,000; original-scaling Muon follows the same multiplier from its 3.0e-02 peak
- The final configured update is positive; update 60,001 is exactly zero
- Official PyTorch Muon with original matrix-shape scaling for hidden matrices; AdamW for embeddings and remaining parameters
Training mixture
- FineWeb-Edu 100BT shuffled: 55%
- Cosmopedia v2: 25%
- FineWeb-HQ: 10%
- FineMath 4+: 10%
FineWeb-Edu supplies the primary educational web text, Cosmopedia supplies synthetic textbook-style coverage, FineWeb-HQ contributes model-filtered, knowledge-rich general web text, and FineMath-4+ supplies mathematical reasoning as ordinary causal-language-model text. The mixture remains fixed for the complete run. There are no benchmark labels or benchmark-specific preprocessing. Dataset revisions are pinned and the same fixed mixture is maintained across both runtime sessions.
Zero-shot evaluation at step 50,000
The four lm-eval tasks use normalized accuracy when supplied by lm-eval
0.4.12, with native bfloat16 weights and float32
likelihood softmax. ArithMark-3 uses its official primary acc_norm metric:
mean continuation-token log likelihood, with context and continuation tokenized
separately after one native BOS prefix. Autocast is not used for evaluation. The four
lm-eval benchmark contexts use the tokenizer's native BOS behavior.
Open SLM Intelligence Index: 6.20 Open SLM Average: 34.68%
| Benchmark | Accuracy |
|---|---|
| HellaSwag | 28.00% |
| ARC-Easy | 33.29% |
| ARC-Challenge | 23.29% |
| PIQA | 56.91% |
ArithMark-3 (acc_norm) |
25.50% |
The Average equally weights HellaSwag, combined ARC (the mean of ARC-Easy and
ARC-Challenge), PIQA, and ArithMark-3. The Intelligence Index first maps random
chance to 0 and perfect accuracy to 100, then weights HellaSwag, combined ARC,
and PIQA at 1.0 and ArithMark-3 at 0.65, matching Open SLM Leaderboard revision
2fbaaa164009c6c0d201ddf245878f5942cc8628.
WikiText-103 validation at this step: loss 3.1840, perplexity 24.14, normalized BPB 1.4397 over 359,037 scored tokens and 1,145,591 normalized UTF-8 bytes, using one initial BOS, 1,024-token windows, and a 512-token stride.
- Downloads last month
- 6
Install from pip and serve model
# Install vLLM from pip: pip install vllm# Start the vLLM server: vllm serve "User01110/cma-mini"# Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "User01110/cma-mini", "prompt": "Once upon a time,", "max_tokens": 512, "temperature": 0.5 }'