File size: 7,284 Bytes
8c36e64
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
144
145
146
147
148
149
150
151
152
153
154
155
156
157
158
---
base_model: MiniMaxAI/MiniMax-M3
library_name: transformers
license: other
license_link: LICENSE
license_name: minimax-community
pipeline_tag: image-text-to-text
tags:
- multimodal
- moe
- agent
- coding
- video
---

<div align="center">
  <img src="https://huggingface.co/buckets/cyankiwi/activation-aware-2.0/resolve/banner/cyankiwi-banner-awq-0.png">
</div>

<div align="left">
  <table align="center" style="border-collapse:collapse; border:none;">
    <tr style="border:none;">
      <td align="right" style="border:none; padding:4px 12px 4px 0;"><b>Version</b></td>
      <td align="left" style="border:none; padding:4px 0;">26.05.01</td>
    </tr>
    <tr style="border:none;">
      <td align="right" style="border:none; padding:4px 12px 4px 0;"><b>Calibration</b></td>
      <td align="left" style="border:none; padding:4px 0;">
      <a href="https://huggingface.co/datasets/cyankiwi/calibration" target="_blank">STEM and Agentic</a>
      </td>
    </tr>
    <tr style="border:none;">
      <td align="right" style="border:none; padding:4px 12px 4px 0;"><b>Languages</b></td>
      <td align="left" style="border:none; padding:4px 0;">
        <code>EN</code> <code>ZH</code> <code>HI</code> <code>AR</code> <code>RU</code>
        <code>JA</code> <code>KO</code> <code>NL</code> <code>FR</code> <code>ES</code>
      </td>
    </tr>
    <tr style="border:none;">
      <td align="right" style="border:none; padding:4px 12px 4px 0;"><b>Model Size</b></td>
      <td align="left" style="border:none; padding:4px 0;">240.30 GB</td>
    </tr>
    <tr style="border:none;">
      <td align="right" style="border:none; padding:4px 12px 4px 0;"><b>Contact</b></td>
      <td align="left" style="border:none; padding:4px 0;">
        <a href="mailto:ton@cyan.kiwi">Email</a>
      </td>
    </tr>
  </table>
</div>

## Serving with vLLM

This checkpoint needs a patched vLLM (MiniMax-M3 compressed-tensors support).
The patch is Python-only, so it installs on top of upstream's **precompiled
binaries** — no CUDA compilation.

### Install

```bash
# uv (skip if already installed)
curl -LsSf https://astral.sh/uv/install.sh | sh

# clone the fork + fetch the upstream base commit
git clone https://github.com/toncao/vllm.git
cd vllm
git remote add upstream https://github.com/vllm-project/vllm.git
git fetch upstream a7fdfeef72323eb3db6f0620e4ea200290d0ca5a
git checkout minimax-m3-compressed-tensors

# Python 3.12 env + install with upstream precompiled kernels
uv venv --python 3.12
source .venv/bin/activate
VLLM_USE_PRECOMPILED=1 uv pip install -e . --torch-backend=auto
```

### Serve

```bash
vllm serve cyankiwi/MiniMax-M3-AWQ-INT4 --block-size 128
```

---

<div align="center">
  <img width="60%" src="figures/logo.svg" alt="MiniMax">
</div>
<hr>

<p align="center">
  <a href="https://agent.minimax.io/" target="_blank"><img src="https://img.shields.io/badge/MiniMax%20Agent-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="MiniMax Agent"></a>
  <a href="https://platform.minimax.io/docs/guides/text-generation" target="_blank"><img src="https://img.shields.io/badge/API-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="API"></a>
  <a href="https://www.minimax.io" target="_blank"><img src="https://img.shields.io/badge/MiniMax%20Website-FF6C37?style=for-the-badge&logo=minimax&logoColor=white" alt="MiniMax Website"></a>
  <br>
  <a href="https://modelscope.cn/organization/minimax" target="_blank" rel="noopener noreferrer"><img alt="ModelScope MiniMax AI" src="https://img.shields.io/badge/ModelScope-MiniMax%20AI-white?labelColor=%23EF3D5D"/></a>
  <a href="https://platform.minimaxi.com/docs/faq/contact-us" target="_blank"><img src="https://img.shields.io/badge/WeChat-07C160?style=for-the-badge&logo=wechat&logoColor=white" alt="WeChat"></a>
  <a href="https://discord.com/invite/DPC4AHFCBw" target="_blank"><img src="https://img.shields.io/badge/Discord-5865F2?style=for-the-badge&logo=discord&logoColor=white" alt="Discord"></a>
  <a href="https://huggingface.co/MiniMaxAI" target="_blank"><img src="https://img.shields.io/badge/Hugging%20Face-FFD21E?style=for-the-badge&logo=huggingface&logoColor=black" alt="Hugging Face"></a>
  <a href="https://github.com/MiniMax-AI/MiniMax-M3" target="_blank"><img src="https://img.shields.io/badge/GitHub-181717?style=for-the-badge&logo=github&logoColor=white" alt="GitHub"></a>
  <a href="https://arxiv.org/abs/2606.13392" target="_blank"><img src="https://img.shields.io/badge/arXiv-2606.13392-B31B1B?style=for-the-badge&logo=arxiv&logoColor=white" alt="arXiv Paper"></a>
  <a href="https://huggingface.co/MiniMaxAI/MiniMax-M3/blob/main/LICENSE" target="_blank"><img src="https://img.shields.io/badge/LICENSE-4CAF50?style=for-the-badge&logo=creativecommons&logoColor=white" alt="LICENSE"></a>
</p>

MiniMax-M3 is a native multimodal model with 1M context. It has ~428B parameters and ~23B activated parameters.

**Highlights:**
- **Native Multimodality:** M3 undergoes mixed-modality training from the very first step, enabling deeper semantic fusion across text, image, and video.
- **Context Scaling via Sparse Attention:** M3 introduces MiniMax Sparse Attention (MSA) to improve long context efficiency. M3 delivers 9× prefill and 15× decode speedups compared to M2 at 1M context, reducing per-token compute to 1/20.
- **Coding & Cowork Capability:** M3 achieves frontier-level performance across long-horizon agentic benchmarks, excelling in both coding and cowork.


<p align="center">
  <img width="100%" src="figures/benchmark.jpeg">
</p>

## MiniMax Sparse Attention (MSA)

M3 is powered by [**MiniMax Sparse Attention (MSA)**](https://github.com/MiniMax-AI/MSA), a high-performance sparse attention operator designed for million-token contexts. Compared with GQA, MSA dramatically reduces the attention compute and memory footprint while preserving model quality.

<p align="center">
  <img width="100%" src="figures/efficiency_gqa_vs_msa.png" alt="GQA vs MSA Efficiency Comparison">
</p>

> 📄 Read the technical report: [arXiv:2606.13392](https://arxiv.org/abs/2606.13392) · [Hugging Face Papers](https://huggingface.co/papers/2606.13392)

## How to Use

- [MiniMax Agent](https://agent.minimax.io/)
- [MiniMax API](https://platform.minimax.io/)

M3 supports two reasoning modes:
- **thinking** — for complex reasoning, agentic tasks, and long-horizon collaboration.
- **non-thinking** — for latency-sensitive scenarios such as chat and code completion.

## Local Deployment

Download the model:

```bash
hf download MiniMaxAI/MiniMax-M3 --local-dir MiniMax-M3
```

We recommend the following inference frameworks (listed alphabetically) to serve the model:

- [SGLang](https://docs.sglang.io/) - see  [SGLang cookbook](https://docs.sglang.io/cookbook/autoregressive/MiniMax/MiniMax-M3).

- [vLLM](https://github.com/vllm-project/vllm) - see [vLLM recipes](https://recipes.vllm.ai/MiniMaxAI/MiniMax-M3).

- [Transformers](https://github.com/huggingface/transformers) - see [Transformers docs](https://huggingface.co/docs/transformers/model_doc/minimax_m3_vl).


### Inference Parameters

We recommend the following parameters for best performance: `temperature=1.0`, `top_p=0.95`, `top_k=40`.

## Contact Us

Contact us at [model@minimax.io](mailto:model@minimax.io).