MOSS-VL-Realtime-NF4 / KV_QUANTIZATION.md
CCCCyx's picture
Add files using upload-large-folder tool
2e600c2 verified
|
Raw
History Blame Contribute Delete
281 Bytes

MOSS-VL quantized KV-cache variant

  • Weights: bitsandbytes NF4 W4A16
  • KV cache: HQQ 8-bit
  • Quantization group size: 64
  • BF16 residual window: 128 tokens per layer
  • Vision/cross-attention weights remain BF16.
  • Requires the quant venv with the selected KV backend installed.