AxionML MiMo-V2.6-Flash-RL-NVFP4

Developed by AxionML for open-source serving and deployment use cases. Part of AxionML's effort to provide ready-to-serve quantized models for the community.

NVFP4 version of XiaomiMiMo/MiMo-V2.6-Flash-RL for Blackwell. The routed experts run as W4A4 NVFP4 on native FP4 Tensor Cores, and their weights are a bit-exact transcode of Xiaomi's released MXFP4 experts: every E2M1 code is kept byte-for-byte and each E8M0 block scale is re-expressed exactly as two E4M3 NVFP4 block scales. All 9,462,349,824 expert scale blocks convert exactly, and none are re-rounded. Every other tensor is copied byte-for-byte from the source checkpoint. The checkpoint is 175 GB (source: 178 GB).

Speed vs Xiaomi's MXFP4 checkpoint

On Blackwell (SM100 family), this checkpoint on the flashinfer_trtllm NVFP4 MoE kernel (W4A4, native FP4 Tensor Cores) is about 1.4x faster than Xiaomi's original MXFP4 checkpoint on --moe-runner-backend flashinfer_mxfp4, the backend that serves the original on B300.

Not yet certain. The ~1.4x is a single measurement by the AxionML team on a separate Blackwell machine, with the branch below. It has not been reproduced on the machine that produced the accuracy numbers on this card, and it will depend on batch size and sequence lengths. A full benchmark table will replace this note.

About NVFP4 quantization: NVFP4 on Blackwell couples a compact E2M1 FP4 codebook with blockwise FP8 (E4M3) scaling over 16-element micro-blocks, so that 4-bit stored values remain numerically useful for neural-network computation. The E2M1 codebook provides a small, nonuniform set of representable magnitudes up to ±6 and relies on saturating behavior rather than IEEE NaN/Inf encodings to maximize usable range per bit. Using an FP8 block scale (rather than power-of-two-only E8M0) enables fractional scales and error-minimizing scale selection. On Blackwell Tensor Cores, native FP4 multipliers exploit E2M1 simplicity to reduce multiplier area while higher-precision FP32 accumulation protects dot-product accuracy.

Ready for commercial and non-commercial use under the MIT License (inherited from the base model).

Quantization Details

Routed experts (47 MoE layers × 256 experts × gate/up/down) NVFP4 W4A4. Weights: exact transcode of the released MXFP4 (E2M1 codes unchanged; per-tensor weight_scale_2 = 2^m, E4M3 block scale 2^(k-m) per 16 elements). Activations: NVFP4 with a static per-tensor global scale of 1.0 (input_scale = 1.0, amax 6 × 448 = 2688)
Fused attention qkv_proj, dense layer-0 MLP, MTP layers FP8 block (128 × 128), unchanged from the source (FP8_PB_WO in the ModelOpt config)
o_proj, embeddings, lm_head, router, vision and audio encoders, DFlash draft BF16, unchanged from the source
KV-cache BF16 (not quantized)
Config ModelOpt MIXED_PRECISION (quantized_layers in config.json / hf_quant_config.json)
Checkpoint size 175 GB (source MXFP4/FP8 checkpoint: 178 GB)
Target hardware Blackwell (verified on 4× B300, sm_103)

Why a unit activation scale. Calibrating expert activations on the dequantized model (agentic-coding, diverse and long-reasoning text plus VQA; ModelOpt max calibration) shows benign expert inputs (amax ≤ 183) but extreme outliers at the down-projection input of the last layers (amax 2,960 at layer 44, 18,560 at layer 46, 311,296 at layer 47). SGLang's NVFP4 MoE kernels use one activation global scale per layer, so a calibrated scale that covers those outliers wastes E4M3 range for every other token. We measured three choices against a BF16-activation reference (same bit-exact weights) on 26K sampled reasoning tokens: raw calibrated scales (+0.0127 NLL/token), calibrated with the down-projection input capped at 2688 (+0.0121), and a unit scale everywhere (+0.0120). They are within noise of each other; we ship the unit scale because it needs no calibration data and measured best.

Usage

Deploy with SGLang

Requires the pinned nightly image plus the mimo-v26-flash-nvfp4-trtllm SGLang branch. The branch makes MiMo-V2's MoE use the correct fused-routing mode (sigmoid + correction-bias top-k, RoutingMethodType.DeepSeekV3); without it the flashinfer_trtllm NVFP4 kernel routes with softmax and generates garbage. The image installs SGLang in editable mode from /sgl-workspace/sglang, so the command swaps that directory for the branch and starts the server on port 30000. Verified on 4× B300:

docker run --gpus all --shm-size=64g --network=host --ipc=host \
  -v ~/.cache/huggingface:/root/.cache/huggingface \
  lmsysorg/sglang:nightly-dev-cu13-20260929-79cafec0 \
  bash -lc '
    cd / && rm -rf /sgl-workspace/sglang &&
    git clone --depth 1 -b mimo-v26-flash-nvfp4-trtllm https://github.com/bzhng-development/sglang.git /sgl-workspace/sglang &&
    sglang serve \
      --trust-remote-code \
      --model-path AxionML/MiMo-V2.6-Flash-RL-NVFP4 \
      --tp 4 \
      --attention-backend fa4 \
      --mm-attention-backend fa4 \
      --moe-runner-backend flashinfer_trtllm \
      --mem-fraction-static 0.85 \
      --reasoning-parser mimo \
      --tool-call-parser mimo \
      --host 0.0.0.0 --port 30000'
  • The leading cd / matters: the image's working directory is /sgl-workspace/sglang, and removing it from inside makes git fail.
  • Attention TP must be 4. The fused qkv_proj is TP=4-interleaved, as in the source checkpoint. For 8 GPUs the SGLang MiMo cookbook uses --tp 8 --dp 2 --enable-dp-attention; we have only verified TP4 with this checkpoint.
  • --attention-backend fa4 is required on Blackwell (asymmetric 192/128 K/V head dims); --mm-attention-backend fa4 for the vision encoder.
  • Pass --moe-runner-backend flashinfer_trtllm explicitly; the default backend crashes on this image ('FusedMoE' object has no attribute 'g1_scale_c').
  • Sampling: Xiaomi recommends temperature=1.0, top_p=0.95. Thinking is on by default; pass chat_template_kwargs={"enable_thinking": false} to turn it off.
  • BF16-activation mode. The same checkpoint also runs with BF16 activations (weight-only FP4, the reference column below): add -e SGLANG_FLASHINFER_CUTEDSL_NVFP4_W4A16=1 to docker run and replace the MoE flags with --quantization modelopt_mixed --moe-runner-backend flashinfer_cutedsl.
from openai import OpenAI

client = OpenAI(base_url="http://localhost:30000/v1", api_key="EMPTY")
r = client.chat.completions.create(
    model="AxionML/MiMo-V2.6-Flash-RL-NVFP4",
    messages=[{"role": "user", "content": "What is 15% of 240?"}],
    temperature=1.0, top_p=0.95, max_tokens=4096,
)
print(r.choices[0].message.reasoning_content)
print(r.choices[0].message.content)

Accuracy

Both columns use this checkpoint's bit-exact expert weights on the same stack (lmsysorg/sglang:v0.5.20-cu130, 4× B300, sgl-eval 0.1.2, thinking on, temperature=1.0, top_p=0.95). The reference runs the experts with BF16 activations (CuTe DSL W4A16 path); the NVFP4 column runs the same W4A4 numerics (bit-exact FP4 weights, NVFP4 activations with input_scale = 1.0) through a different FlashInfer MoE kernel than the flashinfer_trtllm deployment above. Re-measurement on flashinfer_trtllm is pending and will be added here.

Benchmark Budget BF16 activations (reference) NVFP4 W4A4 (this repo)
GSM8K (1319, pass@1) 16K 96.59 96.51
GPQA-Diamond (avg of 4) 32K 76.64 78.28
AIME 2025 (avg of 8) 32K 76.25 78.33
MMMU-Pro (standard, 10 options) 16K 40.23 40.58
Tool call (get_weather, parsed args) 4K pass pass
Image (shape + color identification) 4K pass pass

Share of generations hitting the budget, reference / NVFP4: GPQA 14.9% / 14.3%, AIME 22.5% / 20.8%, MMMU-Pro 17.9% / 17.6%. Differences between the columns are within run-to-run noise (standard error ≈ 1–1.5 points for GPQA, AIME and MMMU-Pro).

Teacher-forced log-likelihood of 26,123 sampled reasoning tokens (48 MMLU-Pro prompts, T=1.0, generated by the BF16-activation reference): mean NLL 0.3377 (reference) vs 0.3493 (NVFP4), top-1 agreement with the reference tokens 87.7% vs 87.4%.

Reproduce (server from Deploy with SGLang on port 30000):

pip install sgl-eval==0.1.2
COMMON="--base-url http://localhost:30000/v1 --temperature 1.0 --top-p 0.95 --chat-template-kwarg enable_thinking=true --num-threads 256"
sgl-eval run gsm8k    $COMMON --max-tokens 16384
sgl-eval run gpqa     $COMMON --max-tokens 32768 --n-repeats 4
sgl-eval run aime25   $COMMON --max-tokens 32768 --n-repeats 8
sgl-eval run mmmu_pro $COMMON --max-tokens 16384

Reproduce the checkpoint

The conversion needs no GPU and no calibration data. tools/build_flash_nvfp4.py reads the released checkpoint, transcodes the MXFP4 experts, copies everything else unchanged, and writes the ModelOpt config:

hf download XiaomiMiMo/MiMo-V2.6-Flash-RL --local-dir ./src
python tools/build_flash_nvfp4.py --src ./src --unit-input-scale --out ./MiMo-V2.6-Flash-RL-NVFP4 --workers 24
# -> expert scale blocks: 9,462,349,824, all-zero: 0, out-of-E4M3-range (re-rounded): 0

The activation-calibration study above used local-inference-lab/quant-toolkit @ 8bdb101 with its MiMo-V2 streaming calibrator; tools/quant-toolkit-8bdb101-mimo-v26-flash.patch adds MXFP4 expert dequantization and ports it to ModelOpt 0.46 / transformers 5.12. It is not needed to rebuild this checkpoint.

Base model

MiMo-V2.6-Flash-RL is the efficiency-balanced checkpoint of Xiaomi's MiMo-V2.6 series: a 309B-total / 15B-active sparse MoE with native text, image, video and audio input and a 1M-token context, trained with one mixed RL run across coding, general agents, visual and cybersecurity tasks. See the original model card and the technical report for architecture, training and evaluation details. This repository changes only the storage format of the routed-expert weights and adds NVFP4 activation quantization for them; the tokenizer, chat template, processor configs, remote code, audio tokenizer and DFlash draft model are copied unchanged.

Limitations

The base model may generate inaccurate, biased or offensive content, and the quantized model inherits these limitations. W4A4 activation quantization changes numerics relative to the released checkpoint even though the expert weights are identical. Please refer to the original model card for full details.

Citation

@misc{mimo2026v26,
  title={MiMo-V2.6: Scaling Reinforcement Learning Towards Self-Improvement},
  author={{Xiaomi MiMo Team}},
  year={2026},
  howpublished={\url{https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL}},
}
Downloads last month
237
Safetensors
Model size
159B params
Tensor type
F32
·
BF16
·
F8_E4M3
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for AxionML/MiMo-V2.6-Flash-RL-NVFP4

Quantized
(35)
this model