Safetensors
nemotron_h

Request access to this model

Please complete the following form.

Log in or Sign Up to review the conditions and access this model content.

Usage

We recommend using vLLM for inference and serving this model.

The following setup is based on the vLLM configuration for NVIDIA Nemotron 3.5 Lightning.

Requirements

  • vLLM 0.27.1+
  • CUDA

Install vLLM

uv venv
source .venv/bin/activate
uv pip install -U vllm --torch-backend auto

vLLM Serve

uv run vllm serve <<model path>> \
  --mamba-backend flashinfer \
  --mamba-cache-mode align \
  --enable-prefix-caching \
  --max-num-batched-tokens 16384 \
  --moe-backend flashinfer_cutlass \
  --tensor-parallel-size 1 \
  --reasoning-parser nemotron_v3 \
  --enable-auto-tool-choice \
  --tool-call-parser qwen3_xml \
  --mamba-ssm-cache-dtype float16 \
  --enable-mamba-cache-stochastic-rounding \
  --mamba-cache-philox-rounds 5 \
  \
  --served-model-name <<model name>> \
  --data-parallel-size 2 \
  --pipeline-parallel-size 1 \
  --async-scheduling \
  --host 0.0.0.0 \
  --port 8000 \
  --max-num-seqs 256

Thinking Mode

We recommend enabling reasoning mode (enable_thinking=True) for improved reasoning performance.

Example using the OpenAI-compatible API:

from openai import OpenAI

client = OpenAI(
    base_url="http://localhost:8000/v1",
    api_key="EMPTY",
)

response = client.chat.completions.create(
    model="<<model name>>",
    messages=[
        {"role": "user", "content": "What is 15 * 37?"}
    ],
    extra_body={
        "chat_template_kwargs": {
            "enable_thinking": True
        }
    },
)

print(response.choices[0].message.content)

Reference

Downloads last month
39
Safetensors
Model size
32B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support