File size: 9,759 Bytes
edebd87
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
159
160
161
162
163
164
165
166
167
168
169
170
171
172
173
174
175
176
177
178
179
180
181
182
183
184
185
186
187
188
189
190
191
192
193
194
195
196
197
198
199
200
201
202
203
204
205
206
207
208
209
210
211
212
213
214
215
216
217
218
219
220
221
222
223
224
225
226
227
228
229
230
231
232
233
234
---
license: apache-2.0
pipeline_tag: text-generation
library_name: transformers
---

# Agnes 2.5 flash base: An Efficient Sparse Mixture-of-Experts Foundation Model

**Agnes 2.5 flash base** is a **202B-parameter sparse Mixture-of-Experts (MoE) base model** with roughly **16B active parameters per token**. It is designed for long-context, high-throughput inference and is released here as an **FP8 checkpoint** that can be served out-of-the-box with `sglang`.

## Introduction

Agnes 2.5 flash base is a decoder-only Transformer that combines several efficiency-oriented components:

1. **Sparse MoE feed-forward layers.** Each of the 48 layers routes every token to **6 of 160 experts** (plus one always-on shared expert). The first 3 layers use deterministic hash routing; the remaining 45 layers use a learned top-k router with auxiliary-loss-free load balancing (`noaux_tc`).
2. **Parallel dense FFN branch.** Layers 3–47 additionally carry a lightweight dense FFN branch (intermediate size 2048) in parallel with the MoE block, increasing per-token capacity at negligible latency cost.
3. **Multi-head Latent Attention (MLA) with KV compression.** Attention uses low-rank query/output projections and a per-layer compressor (compression ratios alternate between 4 and 128 across layers), together with a sparse top-512 token indexer, keeping the KV cache small at very long context.
4. **Hyper-connections.** Residual streams use multi-stream hyper-connections (`hc_mult = 4`) with Sinkhorn-normalized mixing in place of a single residual path.
5. **1M-token context.** YaRN rotary scaling (factor 16 over a 64K base window) extends the usable context to **1,048,576 tokens**.

This repository contains the **base (pre-trained, non-instruction-tuned)** model. It is intended for continued pre-training, fine-tuning, and research; it has not undergone SFT or RLHF, so it should not be expected to follow chat-style instructions reliably.

## Model Zoo


| Model            | Precision | Layers | Experts (active / total) | Params (active / total) | Context   | Hugging Face Model Card |
| ---------------- | --------- | ------ | ------------------------ | ----------------------- | --------- | ----------------------- |
| Agnes 2.5 flash base | FP8       | 48     | 6 + 1 shared / 160       | ~16B / 202B             | 1,048,576 | ✅ this repository       |




### Architecture at a glance


| Hyper-parameter                  | Value                            |
| -------------------------------- | -------------------------------- |
| `hidden_size`                    | 4096                             |
| `num_hidden_layers`              | 48                               |
| `num_attention_heads`            | 64 (`head_dim` 512, RoPE dim 64) |
| `q_lora_rank` / `o_lora_rank`    | 1024 / 1024                      |
| `n_routed_experts`               | 160                              |
| `num_experts_per_tok`            | 6                                |
| `n_shared_experts`               | 1                                |
| `moe_intermediate_size`          | 2048                             |
| `parallel_ffn_intermediate_size` | 2048 (layers 3–47)               |
| `num_hash_layers`                | 3                                |
| `index_topk`                     | 512                              |
| `hc_mult`                        | 4                                |
| `vocab_size`                     | 129,292                          |
| `max_position_embeddings`        | 1,048,576                        |




## Quantization

Weights are stored in **FP8 (e4m3)** with **128×128 block-wise UE8M0 scales** and **dynamic activation quantization**:

```json
"quantization_config": {
  "quant_method": "fp8",
  "fmt": "e4m3",
  "scale_fmt": "ue8m0",
  "weight_block_size": [128, 128],
  "activation_scheme": "dynamic"
}
```

Embeddings, the LM head, normalization layers, router weights and hyper-connection parameters are kept in BF16. Every FP8 linear weight `<name>.weight` is accompanied by a sibling `<name>.scale` tensor (fp32). The checkpoint is ~190 GB across 37 `safetensors` shards.

## Getting Started: Serving with sglang

The recommended way to run Agnes 2.5 flash base is with the **stock** `lmsysorg/sglang:v0.5.16` **Docker image**. Because Agnes support is not yet upstream in sglang, this repository ships the required support files under `[sglang_patch/](./sglang_patch)` together with a launcher script `[serve.sh](./serve.sh)` that overlays them onto the container's sglang package at start-up. **No custom image is needed, and the model directory itself is never modified.**

**Hardware note:** the FP8 checkpoint needs ~190 GB of GPU memory for weights alone. The default configuration uses tensor parallelism over 8 GPUs (e.g. 8× H100/H200 80 GB+).

### 1. Download the model

```shell
pip install -U "huggingface_hub[cli]"
huggingface-cli download <org>/Agnes 2.5 flash base --local-dir ./Agnes 2.5 flash base
```



### 2. Launch the server

```shell
docker run --gpus all --shm-size 64g -p 30001:30002 \
  -v $(pwd)/Agnes 2.5 flash base:/model \
  lmsysorg/sglang:v0.5.16 bash /model/serve.sh
```

`serve.sh` copies `sglang_patch/srt` and `sglang_patch/kernels` into the container's `sglang` package and then execs:

```shell
sglang serve --model-path /model --trust-remote-code --tp 8 \
  --context-length 1048576 --mem-fraction-static 0.90 \
  --host 0.0.0.0 --port 30002
```

Any extra sglang flags can be appended after `serve.sh` and are passed straight through, e.g. a shorter context window to leave more room for the KV cache:

```shell
... lmsysorg/sglang:v0.5.16 bash /model/serve.sh --context-length 262144
```

Model loading takes roughly 10–15 minutes on 8 GPUs. The server is ready once `/health` returns `200`:

```shell
curl http://localhost:30001/health
curl http://localhost:30001/get_model_info
```



### 3. Query the model

Native `/generate` endpoint:

```shell
curl http://localhost:30001/generate \
  -H "Content-Type: application/json" \
  -d '{
    "text": "The three laws of thermodynamics are",
    "sampling_params": {"max_new_tokens": 128, "temperature": 0.7, "top_p": 0.95}
  }'
```

OpenAI-compatible completions endpoint (this is a base model, so prefer `/v1/completions` over `/v1/chat/completions`):

```python
from openai import OpenAI

client = OpenAI(base_url="http://localhost:30001/v1", api_key="EMPTY")
resp = client.completions.create(
    model="default",
    prompt="The three laws of thermodynamics are",
    max_tokens=128,
    temperature=0.7,
    top_p=0.95,
)
print(resp.choices[0].text)
```



### Manual variant (what `serve.sh` does)

If you prefer not to use the launcher script:

```shell
docker run --gpus all --shm-size 64g -p 30001:30002 \
  -v $(pwd)/Agnes 2.5 flash base:/model \
  lmsysorg/sglang:v0.5.16 \
  sh -c "cp -r /model/sglang_patch/srt /model/sglang_patch/kernels \
           /sgl-workspace/sglang/python/sglang/ && \
         exec sglang serve --model-path /model --trust-remote-code --tp 8 \
           --context-length 1048576 --mem-fraction-static 0.90 \
           --host 0.0.0.0 --port 30002"
```

**Important:** the image version must be **exactly** `lmsysorg/sglang:v0.5.16`. The overlay replaces a small set of version-specific files inside sglang; applying it to a different release is not supported.

## Loading with transformers

The repository ships `configuration_agnes.py` and `modeling_agnes.py`, so the model can also be loaded directly with 🤗 transformers using `trust_remote_code=True` (no sglang patch required). Note that the reference PyTorch implementation is intended for inspection, fine-tuning and research rather than high-throughput serving.

```python
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

model_id = "<org>/Agnes 2.5 flash base"
tokenizer = AutoTokenizer.from_pretrained(model_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    model_id,
    trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    device_map="auto",
)

inputs = tokenizer("The three laws of thermodynamics are", return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=64, do_sample=True, temperature=0.7, top_p=0.95)
print(tokenizer.decode(out[0], skip_special_tokens=True))
```



## Repository layout

```
Agnes 2.5 flash base/
├── config.json                     # architecture + FP8 quantization_config
├── generation_config.json
├── configuration_agnes.py          # transformers remote code
├── modeling_agnes.py
├── tokenizer.json / tokenizer_config.json
├── model-000xx-of-00037.safetensors
├── model.safetensors.index.json
├── serve.sh                        # one-command sglang launcher
└── sglang_patch/                   # Agnes support overlay for sglang v0.5.16
    ├── srt/...
    └── kernels/...
```



## Limitations

- **Base model.** No instruction tuning or safety alignment has been applied. Outputs may be incoherent, biased or unsafe; apply your own alignment and filtering before deployment.
- **Memory.** The full FP8 checkpoint requires multi-GPU tensor parallelism; single-GPU inference is not supported.
- **sglang version pin.** The bundled overlay targets sglang `v0.5.16` only.

## License
Both the code repository and the model weights are released under the [Apache License 2.0](LICENSE).


## Citation

If you use Agnes 2.5 flash base in your research, please cite:

```bibtex
@misc{agnes2026flash,
      title={Agnes 2.5 flash base: An Efficient Sparse Mixture-of-Experts Foundation Model},
      author={Agnes AI Team},
      year={2026},
      url={https://huggingface.co/<org>/Agnes 2.5 flash base},
}
```