Instructions to use d1sl1ke/Ling-3.0-flash-oQ4e-mtp with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use d1sl1ke/Ling-3.0-flash-oQ4e-mtp with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir Ling-3.0-flash-oQ4e-mtp d1sl1ke/Ling-3.0-flash-oQ4e-mtp
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
Ling-3.0-flash-oQ4e-mtp
This model was quantized using oQ (oMLX v0.5.8.dev1) mixed-precision quantization.
MTP included but currently is not supported by oMLX.
Original model: Ling-3.0-flash-fp8
This model for M3, M4, M5. For Mac M1, M2 try this.
Benchmarks
Mac M4 Max 40c, 128 GB
| Parameter | Value |
|---|---|
| Model | Ling-3.0-flash-oQ4e-mtp (70.0 GB VRAM) |
| Benchmark Context | Code (Mixed) |
| Generation Length | 128 tokens (fixed) |
Metrics (Single Request)
| Test | TTFT (ms) | TPOT (ms/tok) | pp TPS | tg TPS | E2E Latency | Throughput | Peak Mem |
|---|---|---|---|---|---|---|---|
| pp8192/tg128 | 8,874.6 | 18.14 | 923.1 tok/s | 55.6 tok/s | 11.182 s | 744.1 tok/s | 70.18 GB |
| pp32768/tg128 | 40,822.9 | 19.77 | 802.7 tok/s | 51.0 tok/s | 43.340 s | 759.0 tok/s | 74.59 GB |
| pp65536/tg128 | 99,803.4 | 22.06 | 656.7 tok/s | 45.7 tok/s | 102.614 s | 639.9 tok/s | 80.54 GB |
Quantization details
- Model type: bailing_hybrid
- Bits: 4
- Group size: 64
- Format: MLX safetensors
- Downloads last month
- -
Model size
20B params
Tensor type
BF16
路
U32 路
F32 路
Hardware compatibility
Log In to add your hardware
4-bit
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support
Model tree for d1sl1ke/Ling-3.0-flash-oQ4e-mtp
Base model
inclusionAI/Ling-3.0-flash