# MiniCPM5-1B-NVFP4-W4A4 NVFP4 (W4A4) quantized version of [openbmb/MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B). Quantized with [llm-compressor](https://github.com/vllm-project/llm-compressor) using the `NVFP4` scheme, saved in compressed-tensors format. Both weights and activations are FP4 (W4A4): per-group-16 static fp8 weight scales, and per-group local dynamic fp8 activation scales (static_minmax observer). Calibration: 128 samples drawn from the [`canada-quant/hy3-w4a16-mtp-calibration`](https://huggingface.co/datasets/canada-quant/hy3-w4a16-mtp-calibration) blend (code / math / chat, ~2048 tokens each), formatted through this model's chat template. `lm_head` is left in bf16. Loadable directly in vLLM. Needs a Blackwell GPU with CUDA 13 (e.g. RTX 5090) for native FP4 kernels: ```bash vllm serve MaxDevv/MiniCPM5-1B-NVFP4-W4A4 \ --kv-cache-dtype fp8_e4m3 \ --attention-backend flashinfer ``` --- Original model below: - [openbmb/MiniCPM5-1B](https://huggingface.co/openbmb/MiniCPM5-1B)