Qwen3.8-27B-SwissNeuron-Derisked-GPTQ-INT4

GPTQ W4A16 quantization of SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked.

  • Source revision: repaired direct-answer SFT + fresh latest-SFT DWM α=0.1
  • Format: compressed-tensors GPTQ INT4, group size 128
  • Calibration: frozen SwissNeuron direct-answer corpus rendered with its native chat template
  • Vision tower, MTP draft head, embeddings, and LM head remain BF16
  • Active config retains factor-4 YaRN to 1,048,576 tokens

Runtime

Use a current vLLM stack with Qwen3.8 and compressed-tensors support:

vllm serve SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked-GPTQ-INT4 \
  --trust-remote-code

Thinking modes

Both thinking and non-thinking modes are available through the included Qwen chat template.

Limitations

INT4 is the highest-compression Transformers checkpoint in this release family and may show more degradation than FP8. Validate coding, tool use, multilingual output, and long-context retrieval on your target runtime.

Downloads last month
8
Safetensors
Model size
6B params
Tensor type
BF16
·
I32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SwissNeuron/Qwen3.8-27B-SwissNeuron-Derisked-GPTQ-INT4

Base model

Qwen/Qwen3.8-27B
Quantized
(4)
this model