GPT-OSS-120B modelopt NVFP4 repack

This repository contains an experimental modelopt W4A16_NVFP4 repack of twhitworth/gpt-oss-120b-fp16.

The checkpoint was produced for RTX PRO 6000 / SM120 fallback experiments in the SpecPrefill reproduction workflow. It is intended to be smoke-tested with TensorRT-LLM before benchmark use.

Expected local generation path:

MODEL_ENV_FILE=config/trtllm_modelopt_nvfp4.env \
  bash scripts/trtllm_start_target.sh

MODEL_ENV_FILE=config/trtllm_modelopt_nvfp4.env \
  bash scripts/trtllm_smoke_test_target.sh

Notes:

  • Source checkpoint: twhitworth/gpt-oss-120b-fp16
  • Quantization/repack format: modelopt / W4A16_NVFP4
  • GPT-OSS expert tensor packing is experimental.
  • Verify license and redistribution requirements from the source model before making this repository public.
Downloads last month
22
Safetensors
Model size
66B params
Tensor type
I32
·
BF16
·
F16
·
F8_E4M3
·
U8
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for daeunj/gpt-oss-120b-modelopt-nvfp4

Quantized
(1)
this model