Qwen3.8-27B-SwissNeuron-Derisked-NVFP4

NVIDIA NVFP4 quantization of SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked.

  • Source revision: repaired direct-answer SFT + fresh latest-SFT DWM α=0.1
  • Format: NVIDIA ModelOpt unified Hugging Face checkpoint
  • Language backbone: NVFP4 W4A4 with calibrated block scales
  • Vision tower and MTP draft head remain BF16
  • Intended hardware: NVIDIA Blackwell (RTX 50 series, RTX PRO 6000 Blackwell, B200/B300, GB200/GB300)
  • Active config retains factor-4 YaRN to 1,048,576 tokens

Runtime

Use current vLLM or TensorRT-LLM with ModelOpt NVFP4 support. Blackwell is required for native FP4 Tensor Core execution.

vllm serve SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked-NVFP4 \
  --trust-remote-code

Thinking modes

Both native thinking and non-thinking modes are available through the included Qwen chat template. Set enable_thinking=False only when a direct lower-latency response is desired.

Limitations

NVFP4 is hardware-specific. Do not expect native acceleration on pre-Blackwell GPUs. Re-evaluate critical workloads and long-context retrieval rather than assuming BF16 parity.

Downloads last month
17
Safetensors
Model size
15B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked-NVFP4

Base model

Qwen/Qwen3.8-27B
Quantized
(4)
this model