MOSS-VL-Instruct-0708-NF4 / QUANTIZATION.md
CCCCyx's picture
Add files using upload-large-folder tool
60edb88 verified
|
Raw
History Blame Contribute Delete
424 Bytes
# OpenMOSS-Team/MOSS-VL-Instruct-0708-NF4
- Runtime: Transformers `offline_generate`.
- Weights: bitsandbytes NF4 with double quantization on 240 eligible Linear layers.
- Compute: BF16.
- BF16: first/last four language layers, cross-attention projections, vision encoder/merger, embeddings, norms and `lm_head`.
- KV cache: BF16 (`KV16`); HQQ KV8 is not enabled.
- SGLang compatibility is not claimed for this checkpoint.