WhiteRabbitNeo-V3-7B — MLX (8-bit)

An 8-bit MLX quantization of WhiteRabbitNeo/WhiteRabbitNeo-V3-7B, Kindo's open-weight DevSecOps model, packaged for fast local inference on Apple Silicon via MLX.

At a glance

Base model WhiteRabbitNeo/WhiteRabbitNeo-V3-7B
Base architecture Qwen 2.5 Coder 7B
Parameters ~7.6B
Quantization 8-bit (group size 64, ~8.5 bits/weight)
Format MLX
Prompt format ChatML
Specialization Offensive & defensive cybersecurity / DevSecOps

What this model is

WhiteRabbitNeo V3 is a DevSecOps-focused model from Kindo, fine-tuned from Qwen 2.5 Coder on a large corpus of security Q&A spanning web security, malware analysis, infrastructure-as-code, vulnerability databases, and threat intelligence. It is a minimally-restricted, offensive-capable model designed to assist with security tasks other assistants tend to refuse — vulnerability discovery, exploit reasoning, tooling, and remediation.

Conversion details

A straight quantization of the base weights — no architectural changes, no merges, no fine-tuning.

  • Tool: mlx-lm v0.31.3
  • Command: mlx_lm.convert --hf-path <base> --mlx-path <out> -q --q-bits 8
  • Quantization: 8-bit, group size 64 (~8.5 bpw)

Use with mlx

pip install mlx-lm
from mlx_lm import load, generate

model, tokenizer = load("tthoman79/WhiteRabbitNeo-V3-7B-mlx-8bit")

system = (
    "You are WhiteRabbitNeo, a cybersecurity-expert AI model. "
    "You are an expert in DevOps and Cybersecurity tasks. "
    "Whenever you answer with code, format it with code blocks."
)
messages = [
    {"role": "system", "content": system},
    {"role": "user", "content": "Write a bash script to audit a Linux host for world-writable files."},
]
text = tokenizer.apply_chat_template(messages, add_generation_prompt=True)

response = generate(model, tokenizer, prompt=text, max_tokens=1024, verbose=True)

The model uses the ChatML format and is tuned to operate under a DevSecOps expert system prompt like the one above.

Responsible use & license

This model inherits its license from the base model: Apache 2.0, together with the WhiteRabbitNeo "Extension to Apache-2.0" usage restrictions. Those restrictions prohibit, among other things, using the model or any derivative of it in violation of applicable law or in ways that infringe the rights of others — and they explicitly extend to derivatives, so they apply to this MLX conversion. The original model card holds the complete and authoritative restriction list; review it before use.

This is an offensive-capable security model. Use it only against systems you own or are explicitly authorized to test, for legitimate security research, education, and defense.

Attribution

Model by WhiteRabbitNeo / Kindo, fine-tuned from Qwen 2.5 Coder 7B. This repository provides only an MLX-format quantization; see the original model card for full details.

Downloads last month
22
Safetensors
Model size
2B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for tthoman79/WhiteRabbitNeo-V3-7B-mlx-8bit

Base model

Qwen/Qwen2.5-7B
Quantized
(13)
this model