Agents-A1-ModelOpt-NVFP4

This repository contains an optimized NVFP4 (Group Size 16) quantized checkpoint of nvidia/Agents-A1-FP4 (a 35B parameter hybrid Mamba-Attention Mixture-of-Experts agentic model) produced using NVIDIA ModelOpt and formatted for native vLLM serving.

Model Summary

  • Base Architecture: Agents-A1 MoE (Qwen3_5MoeForCausalLM / hybrid Mamba-Attention)
  • Quantization Method: NVIDIA ModelOpt NVFP4 (W4A16 per-group quantization)
  • Group Size: GS16 (Required configuration for vLLM Marlin FP4 CUDA kernel execution)
  • Total Model Size: ~19.66 GB safetensors weight shards (6 shards)
  • Vocabulary Size: 152,064 tokens (intact)
  • Vision Encoder & MTP: Vision encoder intact (visual.*), vocabulary intact, MTP heads stripped (mtp.*)

Group Size & vLLM Serving Note

Critical Serving Note: While Group Size 128 (GS128) shrinks footprint further, the vLLM Marlin FP4 CUDA kernel (marlin_mm) only supports a group size of 16. Attempts to serve GS128 will result in a engine crash (Invalid thread config). Therefore, GS16 is the only viable serving configuration.

Measured Benchmark Results

  • GSM8K Math Reasoning (20 Representative Samples, temp=0.6, top_p=0.95): Evaluated under local test suite
  • ARC-Challenge Science Reasoning (20 Representative Samples, temp=0.6, top_p=0.95): Evaluated under local test suite

Serving with vLLM

vllm serve Cadododoom/agents-a1-modelopt-nvfp4 \
  --served-model-name nvidia/Agents-A1-FP4 \
  --tensor-parallel-size 2 \
  --quantization compressed-tensors \
  --moe-backend marlin \
  --attention-backend flashinfer \
  --kv-cache-dtype fp8 \
  --max-model-len 112000 \
  --max-num-seqs 4 \
  --max-num-batched-tokens 4096 \
  --gpu-memory-utilization 0.96 \
  --enable-prefix-caching \
  --trust-remote-code \
  --host 0.0.0.0 \
  --port 30000
Downloads last month
32
Safetensors
Model size
18B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support