How to use from
Hermes Agent
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "ZQ-Dev/KAT-Coder-V2.5-Dev-oQ4e-fp16"
Configure Hermes
# Install Hermes:
curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash
hermes setup
# Point Hermes at the local server:
hermes config set model.provider custom
hermes config set model.base_url http://127.0.0.1:8080/v1
hermes config set model.default ZQ-Dev/KAT-Coder-V2.5-Dev-oQ4e-fp16
Run Hermes
hermes
Quick Links

KAT-Coder-V2.5-Dev-oQ4e

A 4-bit enhanced oQ quantization of Kwaipilot/KAT-Coder-V2.5-Dev for Apple Silicon and MLX-compatible runtimes.

Quantization details

The full calibration summary is available in oq_imatrix_report.json.

Note: This quant was created using float16 non-quant weights. Per oMLX, float16 gives ~20% faster prefill on M1/M2 Apple Silicon (native fp16). bfloat16 is safer on M3+ and for numerical stability. For the standard oQe with bfloat16, see this repo.

License

Apache 2.0, inherited from the base model.

Downloads last month
188
Safetensors
Model size
6B params
Tensor type
F16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZQ-Dev/KAT-Coder-V2.5-Dev-oQ4e-fp16

Quantized
(52)
this model

Collection including ZQ-Dev/KAT-Coder-V2.5-Dev-oQ4e-fp16