RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8

This is a quantized version of Qwen/Qwen3.8-2.4T-A95B with MoE layers quantized to NVFP4 and attention layers quantized to FP8 block

Usage

This model is intended for deployment with vLLM. You can serve the model using

vllm serve RedHatAI/Qwen3.8-2.4T-A95B-NVFP4 \
    --data-parallel-size 8 \
    --enable-expert-parallel 8 \
    --reasoning-parser qwen3 \
    --max-num-seqs 140

Creation Process

This model was quantized using LLM Compressor, see https://github.com/vllm-project/llm-compressor/blob/main/docs/key-models/qwen3.5/nvfp4-moe-example.md

Downloads last month
5
Safetensors
Model size
1.4T params
Tensor type
BF16
·
U8
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8

Quantized
(22)
this model