RedHatAI/Qwen3.8-Flash-Next-NVFP4

Model Overview

  • Model Architecture: Qwen4ExpForConditionalGeneration
    • Input: Text / Image / Video
    • Output: Text
  • Model Optimizations:
    • Weight quantization: FP4
    • Activation quantization: FP4
  • Release Date: 2026-08-27
  • Version: 1.0
  • Model Developers: RedHatAI

This model is a quantized version of Qwen/Qwen3.8-Flash-Next. It was evaluated to assess its quality in comparison to the unquantized model.

Model Optimizations

This model was obtained by quantizing the weights and activations of the Mixture-of-Experts (MoE) experts in Qwen/Qwen3.8-Flash-Next to NVFP4 (FP4) data type, ready for inference with vLLM.

This reduces the per-weight precision of the MoE expert parameters from 16 to 4 bits, substantially reducing their memory and disk footprint, while the rest of the model is kept in its original BF16 precision.

Only the weights and activations of the MoE expert linear operators are quantized using LLM Compressor.

Deployment

vLLM Serving

vllm serve RedHatAI/Qwen3.8-Flash-Next-NVFP4 \
  --tensor-parallel-size 4 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_coder \
  --reasoning-parser qwen3

Adjust the tensor-parallel size and other hardware-specific settings to your deployment — see the vLLM recipe for Qwen3.8-Flash-Next.

Creation

This model was created by applying LLM Compressor with calibration samples from open-perfectblend (1024 samples).

Evaluation

On average, the RedHatAI checkpoint shows the highest accuracy recovery of 99.1% over the Inferact's 97.6% and NVIDIA's 98.0%, likely due to differences in calibration data and observer implementations. All of the evals were collected through Inspect with vLLM server and single seed evaluations.

Accuracy

Category Benchmark
Qwen/Qwen3.8-Flash-Next
(baseline)
RedHatAI/Qwen3.8-Flash-Next-NVFP4
(this model)
Inferact/Qwen3.8-Flash-Next-NVFP4
nvidia/Qwen3.8-Flash-Next-NVFP4
Reasoning GPQA Diamond
90.4
92.9
91.4
92.4
Knowledge MMLU-Pro
88.5
88.2
88.0
88.0
Math AIME 2026
96.7
100
100
100
Coding & Agents LiveCodeBench
95.0
95.0
94.0
95.0
Terminal-Bench 2.1
85.5
86.1
84.5
86.4
SWE-bench Verified
80.6
79.4
79.8
79.8
SWE-bench Pro
62.9
63.6
60.9
62.8
DeepSWE 1.1
66.0
62.3
58.1
56.7

Recovery

Recovery is the quantized model's score as a percentage of the unquantized Qwen/Qwen3.8-Flash-Next score, capped at 100%: min(score / unquantized score × 100, 100).

Category Benchmark
RedHatAI/Qwen3.8-Flash-Next-NVFP4
(this model)
Inferact/Qwen3.8-Flash-Next-NVFP4
nvidia/Qwen3.8-Flash-Next-NVFP4
Reasoning GPQA Diamond
100
100
100
Knowledge MMLU-Pro
99.6
99.4
99.4
Math AIME 2026
100
100
100
Coding & Agents LiveCodeBench
100
98.9
100
Terminal-Bench 2.1
100
98.8
100
SWE-bench Verified
98.5
99.0
99.1
SWE-bench Pro
100
96.8
99.9
DeepSWE 1.1
94.3
88.0
85.9
Average
99.1
97.6
98.0
Downloads last month
887
Safetensors
Model size
180B params
Tensor type
BF16
·
U8
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Qwen3.8-Flash-Next-NVFP4

Quantized
(363)
this model

Space using RedHatAI/Qwen3.8-Flash-Next-NVFP4 1