Usage
We recommend using vLLM for inference and serving this model.
The following setup is based on the vLLM configuration for NVIDIA Nemotron 3.5 Lightning.
Requirements
- vLLM 0.27.1+
- CUDA
Install vLLM
uv venv
source .venv/bin/activate
uv pip install -U vllm --torch-backend auto
vLLM Serve
uv run vllm serve <<model path>> \
--mamba-backend flashinfer \
--mamba-cache-mode align \
--enable-prefix-caching \
--max-num-batched-tokens 16384 \
--moe-backend flashinfer_cutlass \
--tensor-parallel-size 1 \
--reasoning-parser nemotron_v3 \
--enable-auto-tool-choice \
--tool-call-parser qwen3_xml \
--mamba-ssm-cache-dtype float16 \
--enable-mamba-cache-stochastic-rounding \
--mamba-cache-philox-rounds 5 \
\
--served-model-name <<model name>> \
--data-parallel-size 2 \
--pipeline-parallel-size 1 \
--async-scheduling \
--host 0.0.0.0 \
--port 8000 \
--max-num-seqs 256
Thinking Mode
We recommend enabling reasoning mode (enable_thinking=True) for improved reasoning performance.
Example using the OpenAI-compatible API:
from openai import OpenAI
client = OpenAI(
base_url="http://localhost:8000/v1",
api_key="EMPTY",
)
response = client.chat.completions.create(
model="<<model name>>",
messages=[
{"role": "user", "content": "What is 15 * 37?"}
],
extra_body={
"chat_template_kwargs": {
"enable_thinking": True
}
},
)
print(response.choices[0].message.content)
Reference
- Downloads last month
- 39
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support