Qwythos-9B-v2-MLX-oQ8-mtp

This model was quantized using oQ (oMLX v0.5.4.dev1) mixed-precision quantization.

Quantization details

  • Model type: qwen3_5
  • Bits: 8
  • Group size: 64
  • Format: MLX safetensors

Testing

I used these params recently for a complex reasoning task requring RAG and high-level thinking. Results were slow but exceptionally strong.

Added kwargs forced reasoning effort = Max

ACTIVE MODEL

Qwythos-9B-v2-MLX-oQ8-mtp

TEMPERATURE

0.6

MAX TOKENS

16384

MIN P

0.95 TOP K 20

REP. PENALTY

1.05

PRESENCE PENALTY

Default

THINKING

On (Unlimited)

Downloads last month
309
Safetensors
Model size
3B params
Tensor type
BF16
·
U32
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for brainworkup/Qwythos-9B-v2-MLX-oQ8-mtp

Finetuned
Qwen/Qwen3.5-9B
Quantized
(29)
this model