MOSS-VL-Instruct-0708-NF4 / QUANTIZATION.md
CCCCyx's picture
Add files using upload-large-folder tool
60edb88 verified
|
Raw
History Blame Contribute Delete
424 Bytes

OpenMOSS-Team/MOSS-VL-Instruct-0708-NF4

  • Runtime: Transformers offline_generate.
  • Weights: bitsandbytes NF4 with double quantization on 240 eligible Linear layers.
  • Compute: BF16.
  • BF16: first/last four language layers, cross-attention projections, vision encoder/merger, embeddings, norms and lm_head.
  • KV cache: BF16 (KV16); HQQ KV8 is not enabled.
  • SGLang compatibility is not claimed for this checkpoint.