Instructions to use AxionML/MiMo-V2.6-Flash-RL-NVFP4 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use AxionML/MiMo-V2.6-Flash-RL-NVFP4 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="AxionML/MiMo-V2.6-Flash-RL-NVFP4", trust_remote_code=True) messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("AxionML/MiMo-V2.6-Flash-RL-NVFP4", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use AxionML/MiMo-V2.6-Flash-RL-NVFP4 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "AxionML/MiMo-V2.6-Flash-RL-NVFP4" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/MiMo-V2.6-Flash-RL-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/AxionML/MiMo-V2.6-Flash-RL-NVFP4
- SGLang
How to use AxionML/MiMo-V2.6-Flash-RL-NVFP4 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "AxionML/MiMo-V2.6-Flash-RL-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/MiMo-V2.6-Flash-RL-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "AxionML/MiMo-V2.6-Flash-RL-NVFP4" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "AxionML/MiMo-V2.6-Flash-RL-NVFP4", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use AxionML/MiMo-V2.6-Flash-RL-NVFP4 with Docker Model Runner:
docker model run hf.co/AxionML/MiMo-V2.6-Flash-RL-NVFP4
AxionML MiMo-V2.6-Flash-RL-NVFP4
Developed by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.
NVFP4 version of XiaomiMiMo/MiMo-V2.6-Flash-RL for Blackwell. The routed experts run as W4A4 NVFP4 on native FP4 Tensor Cores, and their weights are a bit-exact transcode of Xiaomi's released MXFP4 experts: every E2M1 code is kept byte-for-byte and each E8M0 block scale is re-expressed exactly as two E4M3 NVFP4 block scales. All 9,462,349,824 expert scale blocks convert exactly, and none are re-rounded. Every other tensor is copied byte-for-byte from the source checkpoint. The checkpoint is 175 GB (source: 178 GB).
Speed vs Xiaomi's MXFP4 checkpoint
On Blackwell (SM100 family), this checkpoint on the flashinfer_trtllm NVFP4 MoE kernel (W4A4, native FP4 Tensor Cores) is about 1.4x faster than Xiaomi's original MXFP4 checkpoint on --moe-runner-backend flashinfer_mxfp4, the backend that serves the original on B300.
Not yet certain. The ~1.4x is a single measurement by the AxionML team on a separate Blackwell machine, with the branch below. It has not been reproduced on the machine that produced the accuracy numbers on this card, and it will depend on batch size and sequence lengths. A full benchmark table will replace this note.
About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity to reduce multiplier area while higher-precision FP32 accumulation protects dot-product accuracy.
Ready for commercial and non-commercial use under the MIT License (inherited from the base model).
Quantization Details
| Routed experts (47 MoE layers × 256 experts × gate/up/down) | NVFP4 W4A4. Weights: exact transcode of the released MXFP4 (E2M1 codes unchanged; per-tensor weight_scale_2 = 2^m, E4M3 block scale 2^(k-m) per 16 elements). Activations: NVFP4 with a static per-tensor global scale of 1.0 (input_scale = 1.0, amax 6 × 448 = 2688) |
Fused attention qkv_proj, dense layer-0 MLP, MTP layers |
FP8 block (128 × 128), unchanged from the source (FP8_PB_WO in the ModelOpt config) |
o_proj, embeddings, lm_head, router, vision and audio encoders, DFlash draft |
BF16, unchanged from the source |
| KV-cache | BF16 (not quantized) |
| Config | ModelOpt MIXED_PRECISION (quantized_layers in config.json / hf_quant_config.json) |
| Checkpoint size | 175 GB (source MXFP4/FP8 checkpoint: 178 GB) |
| Target hardware | Blackwell (verified on 4× B300, sm_103) |
Why a unit activation scale. Calibrating expert activations on the dequantized model (agentic-coding, diverse and long-reasoning text plus VQA; ModelOpt max calibration) shows benign expert inputs (amax ≤ 183) but extreme outliers at the down-projection input of the last layers (amax 2,960 at layer 44, 18,560 at layer 46, 311,296 at layer 47). SGLang's NVFP4 MoE kernels use one activation global scale per layer, so a calibrated scale that covers those outliers wastes E4M3 range for every other token. We measured three choices against a BF16-activation reference (same bit-exact weights) on 26K sampled reasoning tokens: raw calibrated scales (+0.0127 NLL/token), calibrated with the down-projection input capped at 2688 (+0.0121), and a unit scale everywhere (+0.0120). They are within noise of each other; we ship the unit scale because it needs no calibration data and measured best.
Usage
Deploy with SGLang
Requires the pinned nightly image plus the mimo-v26-flash-nvfp4-trtllm SGLang branch. The branch makes MiMo-V2's MoE use the correct fused-routing mode (sigmoid + correction-bias top-k, RoutingMethodType.DeepSeekV3); without it the flashinfer_trtllm NVFP4 kernel routes with softmax and generates garbage. The image installs SGLang in editable mode from /sgl-workspace/sglang, so the command swaps that directory for the branch and starts the server on port 30000. Verified on 4× B300:
docker run --gpus all --shm-size=64g --network=host --ipc=host \
-v ~/.cache/huggingface:/root/.cache/huggingface \
lmsysorg/sglang:nightly-dev-cu13-20260929-79cafec0 \
bash -lc '
cd / && rm -rf /sgl-workspace/sglang &&
git clone --depth 1 -b mimo-v26-flash-nvfp4-trtllm https://github.com/bzhng-development/sglang.git /sgl-workspace/sglang &&
sglang serve \
--trust-remote-code \
--model-path AxionML/MiMo-V2.6-Flash-RL-NVFP4 \
--tp 4 \
--attention-backend fa4 \
--mm-attention-backend fa4 \
--moe-runner-backend flashinfer_trtllm \
--mem-fraction-static 0.85 \
--reasoning-parser mimo \
--tool-call-parser mimo \
--host 0.0.0.0 --port 30000'
- The leading
cd /matters: the image's working directory is/sgl-workspace/sglang, and removing it from inside makesgitfail. - Attention TP must be 4. The fused
qkv_projis TP=4-interleaved, as in the source checkpoint. For 8 GPUs the SGLang MiMo cookbook uses--tp 8 --dp 2 --enable-dp-attention; we have only verified TP4 with this checkpoint. --attention-backend fa4is required on Blackwell (asymmetric 192/128 K/V head dims);--mm-attention-backend fa4for the vision encoder.- Pass
--moe-runner-backend flashinfer_trtllmexplicitly; the default backend crashes on this image ('FusedMoE' object has no attribute 'g1_scale_c'). - Sampling: Xiaomi recommends
temperature=1.0, top_p=0.95. Thinking is on by default; passchat_template_kwargs={"enable_thinking": false}to turn it off. - BF16-activation mode. The same checkpoint also runs with BF16 activations (weight-only FP4, the reference column below): add
-e SGLANG_FLASHINFER_CUTEDSL_NVFP4_W4A16=1todocker runand replace the MoE flags with--quantization modelopt_mixed --moe-runner-backend flashinfer_cutedsl.
from openai import OpenAI
client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
r = client.chat.completions.create(
model="AxionML/MiMo-V2.6-Flash-RL-NVFP4",
messages=[{"role": "user", "content": "What is 15% of 240?"}],
temperature=1.0, top_p=0.95, max_tokens=4096,
)
print(r.choices[0].message.reasoning_content)
print(r.choices[0].message.content)
Accuracy
Both columns use this checkpoint's bit-exact expert weights on the same stack (lmsysorg/sglang:v0.5.20-cu130, 4× B300, sgl-eval 0.1.2, thinking on, temperature=1.0, top_p=0.95). The reference runs the experts with BF16 activations (CuTe DSL W4A16 path); the NVFP4 column runs the same W4A4 numerics (bit-exact FP4 weights, NVFP4 activations with input_scale = 1.0) through a different FlashInfer MoE kernel than the flashinfer_trtllm deployment above. Re-measurement on flashinfer_trtllm is pending and will be added here.
| Benchmark | Budget | BF16 activations (reference) | NVFP4 W4A4 (this repo) |
|---|---|---|---|
| GSM8K (1319, pass@1) | 16K | 96.59 | 96.51 |
| GPQA-Diamond (avg of 4) | 32K | 76.64 | 78.28 |
| AIME 2025 (avg of 8) | 32K | 76.25 | 78.33 |
| MMMU-Pro (standard, 10 options) | 16K | 40.23 | 40.58 |
| Tool call (get_weather, parsed args) | 4K | pass | pass |
| Image (shape + color identification) | 4K | pass | pass |
Share of generations hitting the budget, reference / NVFP4: GPQA 14.9% / 14.3%, AIME 22.5% / 20.8%, MMMU-Pro 17.9% / 17.6%. Differences between the columns are within run-to-run noise (standard error ≈ 1–1.5 points for GPQA, AIME and MMMU-Pro).
Teacher-forced log-likelihood of 26,123 sampled reasoning tokens (48 MMLU-Pro prompts, T=1.0, generated by the BF16-activation reference): mean NLL 0.3377 (reference) vs 0.3493 (NVFP4), top-1 agreement with the reference tokens 87.7% vs 87.4%.
Reproduce (server from Deploy with SGLang on port 30000):
pip install sgl-eval==0.1.2
COMMON="--base-url http://localhost:30000/v1 --temperature 1.0 --top-p 0.95 --chat-template-kwarg enable_thinking=true --num-threads 256"
sgl-eval run gsm8k $COMMON --max-tokens 16384
sgl-eval run gpqa $COMMON --max-tokens 32768 --n-repeats 4
sgl-eval run aime25 $COMMON --max-tokens 32768 --n-repeats 8
sgl-eval run mmmu_pro $COMMON --max-tokens 16384
Reproduce the checkpoint
The conversion needs no GPU and no calibration data. tools/build_flash_nvfp4.py reads the released checkpoint, transcodes the MXFP4 experts, copies everything else unchanged, and writes the ModelOpt config:
hf download XiaomiMiMo/MiMo-V2.6-Flash-RL --local-dir ./src
python tools/build_flash_nvfp4.py --src ./src --unit-input-scale --out ./MiMo-V2.6-Flash-RL-NVFP4 --workers 24
# -> expert scale blocks: 9,462,349,824, all-zero: 0, out-of-E4M3-range (re-rounded): 0
The activation-calibration study above used local-inference-lab/quant-toolkit @ 8bdb101 with its MiMo-V2 streaming calibrator; tools/quant-toolkit-8bdb101-mimo-v26-flash.patch adds MXFP4 expert dequantization and ports it to ModelOpt 0.46 / transformers 5.12. It is not needed to rebuild this checkpoint.
Base model
MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of Xiaomi's MiMo-V2.6 series: a 309B-total / 15B-active sparse MoE with native text, image, video and audio input and a 1M-token context, trained with one mixed RL run across coding, general agents, visual and cybersecurity tasks. See the original model card and the technical report for architecture, training and evaluation details. This repository changes only the storage format of the routed-expert weights and adds NVFP4 activation quantization for them; the tokenizer, chat template, processor configs, remote code, audio tokenizer and DFlash draft model are copied unchanged.
Limitations
The base model may generate inaccurate, biased or offensive content, and the quantized model inherits these limitations. W4A4 activation quantization changes numerics relative to the released checkpoint even though the expert weights are identical. Please refer to the original model card for full details.
Citation
@misc{mimo2026v26,
title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
author={{Xiaomi MiMo Team}},
year={2026},
howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}
- Downloads last month
- 237
Model tree for AxionML/MiMo-V2.6-Flash-RL-NVFP4
Base model
XiaomiMiMo/MiMo-V2.6-Flash-RL