| --- |
| license: apache-2.0 |
| base_model: nvidia/Agents-A1-FP4 |
| tags: |
| - modelopt |
| - nvfp4 |
| - fp4 |
| - gs16 |
| - compressed-tensors |
| - vllm |
| - moe |
| - mamba |
| - tensor-parallel |
| --- |
| |
| # Agents-A1-ModelOpt-NVFP4 |
|
|
| This repository contains an optimized **NVFP4 (Group Size 16)** quantized checkpoint of `nvidia/Agents-A1-FP4` (a 35B parameter hybrid Mamba-Attention Mixture-of-Experts agentic model) produced using **NVIDIA ModelOpt** and formatted for native **vLLM** serving. |
|
|
| ## Model Summary |
| * **Base Architecture**: Agents-A1 MoE (`Qwen3_5MoeForCausalLM` / hybrid Mamba-Attention) |
| * **Quantization Method**: NVIDIA ModelOpt NVFP4 (W4A16 per-group quantization) |
| * **Group Size**: **GS16** (Required configuration for vLLM Marlin FP4 CUDA kernel execution) |
| * **Total Model Size**: ~19.66 GB safetensors weight shards (6 shards) |
| * **Vocabulary Size**: 152,064 tokens (intact) |
| * **Vision Encoder & MTP**: Vision encoder intact (`visual.*`), vocabulary intact, MTP heads stripped (`mtp.*`) |
|
|
| ## Group Size & vLLM Serving Note |
| > **Critical Serving Note**: While Group Size 128 (GS128) shrinks footprint further, the vLLM Marlin FP4 CUDA kernel (`marlin_mm`) only supports a group size of 16. Attempts to serve GS128 will result in a engine crash (`Invalid thread config`). Therefore, **GS16 is the only viable serving configuration**. |
| |
| ## Measured Benchmark Results |
| * **GSM8K Math Reasoning (20 Representative Samples, temp=0.6, top_p=0.95)**: Evaluated under local test suite |
| * **ARC-Challenge Science Reasoning (20 Representative Samples, temp=0.6, top_p=0.95)**: Evaluated under local test suite |
| |
| ## Serving with vLLM |
| |
| ```bash |
| vllm serve Cadododoom/agents-a1-modelopt-nvfp4 \ |
| --served-model-name nvidia/Agents-A1-FP4 \ |
| --tensor-parallel-size 2 \ |
| --quantization compressed-tensors \ |
| --moe-backend marlin \ |
| --attention-backend flashinfer \ |
| --kv-cache-dtype fp8 \ |
| --max-model-len 112000 \ |
| --max-num-seqs 4 \ |
| --max-num-batched-tokens 4096 \ |
| --gpu-memory-utilization 0.96 \ |
| --enable-prefix-caching \ |
| --trust-remote-code \ |
| --host 0.0.0.0 \ |
| --port 30000 |
| ``` |
| |