KAT-Coder-V2.5-Dev-oQ8e

An 8-bit enhanced oQ quantization of Kwaipilot/KAT-Coder-V2.5-Dev for Apple Silicon and MLX-compatible runtimes.

Quantization details

The full calibration summary is available in oq_imatrix_report.json.

Note: This quant was created using float16 non-quant weights. Per oMLX, float16 gives ~20% faster prefill on M1/M2 Apple Silicon (native fp16). bfloat16 is safer on M3+ and for numerical stability. For the standard oQe with bfloat16, see this repo.

License

Apache 2.0, inherited from the base model.

Downloads last month
-
Safetensors
Model size
10B params
Tensor type
F16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZQ-Dev/KAT-Coder-V2.5-Dev-oQ8e-fp16

Quantized
(50)
this model

Collection including ZQ-Dev/KAT-Coder-V2.5-Dev-oQ8e-fp16