viraat57's picture
Upload README.md with huggingface_hub
e7a32ba verified
|
Raw
History Blame Contribute Delete
4.92 kB
---
pipeline_tag: text-generation
base_model: deepseek-ai/DeepSeek-R1
license: mit
library_name: modelopt
tags:
- ModelOpt
- DeepSeekR1
- quantized
- FP4
- MLPerf
---
# Model Overview
## Description:
This model is a quantized version of DeepSeek AI's DeepSeek R1 model, produced with the official [NVIDIA TensorRT Model Optimizer DeepSeek recipe](https://github.com/NVIDIA/Model-Optimizer/tree/089c06e4/examples/deepseek) and calibrated with the [MLPerf Inference deepseek-r1 calibration dataset](https://github.com/mlcommons/inference/blob/master/language/deepseek-r1/README.md). It is quantized with FP4 weights and activations (the same quantization scheme as [nvidia/DeepSeek-R1-FP4-v2](https://huggingface.co/nvidia/DeepSeek-R1-FP4-v2), including the `wo` module in attention layers) and is ready for inference with TensorRT-LLM. This model is ready for commercial/non-commercial use.
## Third-Party Community Consideration
This model is not owned or developed by NVIDIA or DeepSeek AI. It has been produced by a third party from the open DeepSeek R1 checkpoint for the MLPerf Inference deepseek-r1 benchmark. See the original [DeepSeek R1 Model Card](https://huggingface.co/deepseek-ai/DeepSeek-R1).
### License/Terms of Use:
[MIT](https://huggingface.co/datasets/choosealicense/licenses/blob/main/markdown/mit.md)
## Model Architecture:
**Architecture Type:** Transformers
**Network Architecture:** DeepSeek R1
## Input:
**Input Type(s):** Text
**Input Format(s):** String
**Input Parameters:** 1D (One Dimensional): Sequences
**Other Properties Related to Input:** Context length up to 128K. Per DeepSeek's usage recommendations: temperature 0.5–0.7 (0.6 recommended), no system prompt, and for math problems ask for step-by-step reasoning with the final answer in `\boxed{}`.
## Output:
**Output Type(s):** Text
**Output Format:** String
**Output Parameters:** 1D (One Dimensional): Sequences
## Software Integration:
**Supported Runtime Engine(s):**
- TensorRT-LLM
**Supported Hardware Microarchitecture Compatibility:**
- NVIDIA Blackwell
**Preferred Operating System(s):**
- Linux
## Model Version(s):
Quantized with [NVIDIA TensorRT Model Optimizer](https://github.com/NVIDIA/Model-Optimizer), main branch @ [`089c06e4`](https://github.com/NVIDIA/Model-Optimizer/tree/089c06e4) (July 2026), using the [`examples/deepseek`](https://github.com/NVIDIA/Model-Optimizer/tree/089c06e4/examples/deepseek) recipe (`deepseek_v3/ptq.py` for calibration, `deepseek_v3/quantize_fp8_to_nvfp4.sh` for weight conversion).
## Datasets:
**Calibration Dataset:** `mlperf_deepseek_r1_calibration_dataset_500_fp8_eval` — the official 500-sample MLPerf Inference deepseek-r1 calibration set. See the [MLPerf deepseek-r1 reference README](https://github.com/mlcommons/inference/blob/master/language/deepseek-r1/README.md) for dataset details and download instructions.
- Data collection method: Hybrid: Human, Automated
- Labeling method: Hybrid: Human, Automated
## Inference:
**Engine:** TensorRT-LLM
**Test Hardware:** NVIDIA GB200 NVL72
## Post Training Quantization
This model was obtained by quantizing the weights and activations of DeepSeek R1 to FP4 data type, ready for inference with TensorRT-LLM. Only the weights and activations of the linear operators within the transformer blocks are quantized (fine-grained FP4 block scaling with group size 16; the KV cache is quantized to FP8; MLA attention projections remain in higher precision). This optimization reduces the number of bits per parameter from 8 to 4, reducing the disk size and GPU memory requirements by approximately 1.6x.
Calibration runs the official recipe unmodified except for the calibration dataloader, which reads the MLPerf deepseek-r1 calibration set (500 samples, max sequence length 2048, batch size 4).
## Usage
### Deploy with TensorRT-LLM
To deploy the quantized checkpoint with [TensorRT-LLM](https://github.com/NVIDIA/TensorRT-LLM) LLM API, use the following sample code (requires 8 GPUs of Blackwell generation or later, and TensorRT-LLM built from source with the latest main branch):
```python
from tensorrt_llm import SamplingParams
from tensorrt_llm._torch import LLM
def main():
prompts = [
"Hello, my name is",
"The president of the United States is",
"The capital of France is",
"The future of AI is",
]
sampling_params = SamplingParams(max_tokens=32)
llm = LLM(model="centml/DeepSeek-R1-NVFP4-v2-mlpinf", tensor_parallel_size=8, enable_attention_dp=True)
outputs = llm.generate(prompts, sampling_params)
# Print the outputs.
for output in outputs:
prompt = output.prompt
generated_text = output.outputs[0].text
print(f"Prompt: {prompt!r}, Generated text: {generated_text!r}")
# The entry point of the program need to be protected for spawning processes.
if __name__ == '__main__':
main()
```