MOSS-VL-Realtime-NF4 / KV_QUANTIZATION.md
CCCCyx's picture
Add files using upload-large-folder tool
2e600c2 verified
|
Raw
History Blame Contribute Delete
281 Bytes
# MOSS-VL quantized KV-cache variant
- Weights: bitsandbytes NF4 W4A16
- KV cache: HQQ 8-bit
- Quantization group size: 64
- BF16 residual window: 128 tokens per layer
- Vision/cross-attention weights remain BF16.
- Requires the quant venv with the selected KV backend installed.