How to use from
Pi
Start the MLX server
# Install MLX LM:
uv tool install mlx-lm
# Start a local OpenAI-compatible server:
mlx_lm.server --model "ZQ-Dev/KAT-Coder-V2.5-Dev-oQ2e-fp16"
Configure the model in Pi
# Install Pi:
npm install -g @mariozechner/pi-coding-agent
# Add to ~/.pi/agent/models.json:
{
  "providers": {
    "mlx-lm": {
      "baseUrl": "http://localhost:8080/v1",
      "api": "openai-completions",
      "apiKey": "none",
      "models": [
        {
          "id": "ZQ-Dev/KAT-Coder-V2.5-Dev-oQ2e-fp16"
        }
      ]
    }
  }
}
Run Pi
# Start Pi in your project directory:
pi
Quick Links

KAT-Coder-V2.5-Dev-oQ2e

A 2-bit enhanced oQ quantization of Kwaipilot/KAT-Coder-V2.5-Dev for Apple Silicon and MLX-compatible runtimes.

Quantization details

The full calibration summary is available in oq_imatrix_report.json.

Note: This quant was created using float16 non-quant weights. Per oMLX, float16 gives ~20% faster prefill on M1/M2 Apple Silicon (native fp16). bfloat16 is safer on M3+ and for numerical stability. For the standard oQe with bfloat16, see this repo.

License

Apache 2.0, inherited from the base model.

Downloads last month
211
Safetensors
Model size
4B params
Tensor type
F16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ZQ-Dev/KAT-Coder-V2.5-Dev-oQ2e-fp16

Quantized
(52)
this model

Collection including ZQ-Dev/KAT-Coder-V2.5-Dev-oQ2e-fp16