Ling-3.0-flash-oQ4e-mtp

This model was quantized using oQ (oMLX v0.5.8.dev1) mixed-precision quantization.

MTP included but currently is not supported by oMLX.

Original model: Ling-3.0-flash-fp8

This model for M3, M4, M5. For Mac M1, M2 try this.

Benchmarks

Mac M4 Max 40c, 128 GB

Parameter Value
Model Ling-3.0-flash-oQ4e-mtp (70.0 GB VRAM)
Benchmark Context Code (Mixed)
Generation Length 128 tokens (fixed)

Metrics (Single Request)

Test TTFT (ms) TPOT (ms/tok) pp TPS tg TPS E2E Latency Throughput Peak Mem
pp8192/tg128 8,874.6 18.14 923.1 tok/s 55.6 tok/s 11.182 s 744.1 tok/s 70.18 GB
pp32768/tg128 40,822.9 19.77 802.7 tok/s 51.0 tok/s 43.340 s 759.0 tok/s 74.59 GB
pp65536/tg128 99,803.4 22.06 656.7 tok/s 45.7 tok/s 102.614 s 639.9 tok/s 80.54 GB

Quantization details

  • Model type: bailing_hybrid
  • Bits: 4
  • Group size: 64
  • Format: MLX safetensors
Downloads last month
-
Safetensors
Model size
20B params
Tensor type
BF16
U32
F32
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for d1sl1ke/Ling-3.0-flash-oQ4e-mtp

Quantized
(17)
this model