TheWirelessPhoenix's picture
Upload README.md with huggingface_hub
36cd38a verified
|
Raw
History Blame Contribute Delete
2.95 kB
---
license: mit
library_name: llama.cpp
tags:
- gguf
- llama.cpp
- bailing-hybrid
- ling
- hybrid-kda-mla
---
# Ling 3.0 Tiny - experimental GGUF conversions
This repository contains highly experimental GGUF conversions of [`inclusionAI/Ling-3.0-tiny`](https://huggingface.co/inclusionAI/Ling-3.0-tiny) for an experimental `bailing-hybrid` implementation in llama.cpp.
These files should **not** be treated as stable, production-ready, or numerically certified conversions. The hybrid architecture combines KDA recurrent layers with gated MLA layers and MoE layers, and requires the accompanying experimental llama.cpp source snapshot. A stock llama.cpp checkout may not recognize this architecture.
## Files
| File | Format | Size | SHA-256 |
|---|---:|---:|---|
| `Ling-3.0-tiny-F16.gguf` | F16 reference conversion | 15,803,475,488 bytes | `51b812d28fdab4d13caf9c5325fa1c99aeb2687644ebcf2bef393d9be51c82dc` |
| `Ling-3.0-tiny-Q5_K_M.gguf` | Q5_K_M | 5,635,443,552 bytes | `0f8159fa1a72d1997f89c121f9bedcf502ad1e8fc2f6e6196b22818fc3378080` |
| `Ling-3.0-tiny-Q4_K_M.gguf` | Q4_K_M | 4,823,894,880 bytes | `b1cffbbb88770fe3d0de375892f540c5a8997fa441c5ea0a94b19319dcb9a55c` |
The quantized files were created directly from the F16 GGUF with llama.cpp's `llama-quantize`. The standard mixed K-quant behavior was used; six small tensors in each quantization required the quantizer's fallback type because their shapes are not compatible with the requested block format.
## Validation performed
- GGUF metadata and tensor descriptors loaded successfully.
- The F16 conversion contains 526 tensors and records the `bailing-hybrid` architecture.
- Q5_K_M and Q4_K_M both loaded successfully in the experimental llama.cpp runtime.
- On an Apple M5, both quantized files offloaded all 25 model layers to `MTL0` and completed a short generation test.
- The build used for validation had `GGML_METAL=ON` with embedded Metal shaders.
Metal support belongs to the llama.cpp runtime rather than the GGUF file itself. Use a build with the experimental architecture changes and Metal enabled if you want GPU offload on Apple hardware.
## Important limitations
This is a **very, very experimental** port. It has not received full long-context, concurrency, state rollback, cross-backend, or exact Hugging Face logits-parity validation. The successful smoke tests are not a guarantee of general correctness or model quality. The F16 file is a reference conversion, not a claim of upstream compatibility.
If stability, reproducibility, or production reliability is important, please use a more established model conversion and a llama.cpp release with official support instead of these files.
The `Ling3.0-tiny-llama.cpp-source-experimental.zip` archive contains the source snapshot used for this conversion, including the local experimental changes. It intentionally excludes `.git`, `.venv`, and generated `build` contents, including `build/bin`.