Soulfate24's picture
## [1.3.1] - 2026-08-20
ae55163 verified
|
Raw
History Blame Contribute Delete
10 kB
---
license: apache-2.0
tags:
- quantization
- gguf
- autoround
- imatrix
- hybrid-quantization
- llama.cpp
- text-generation
---
# AutoRound + ASHQ1 Double-Quantization Suite
The **AutoRound + ASHQ1 Suite** delivers a complete pipeline for creating ultra-high-fidelity GGUF models. By combining gradient-guided weight reorganization (**AutoRound W4A16**) with fine-grained activation-aware tensor assignment (**ASHQ1 Imatrix Engine**), this suite establishes a new standard for low-bit LLM compression.
---
## 🌟 Key Architecture & Highlights
```
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Safetensors (Raw / BF16)β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚ 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ AutoRound Optimization β”‚ ──► Iterative sign-rounding & Hessian estimation
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚ Streaming dequantization + GGUF encapsulation
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ AutoRound-Infused BF16 β”‚ ──► Lineage recorded in sidecar metadata
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚ 01_create-calibration-dataset-and-imatrix.py
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Multi-Source Imatrix β”‚ ──► Agentic, Frontier, Logic & Diversity corpus
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
β”‚ 02_BF16-GGUF-to-ASHQ1.py (ASHQ1 Engine)
β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Standardized ASHQ1 Tiers: Nano β€’ Mini β€’ Compact β€’ Quality β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
```
### 1. Dual-Phase Quantization Synergy
* **Phase 1 (AutoRound W4A16)**: Reconditions original full-precision matrices via small-sample Hessian compensation. The optimized rounding directions remain intact when converted into BF16 GGUF containers.
* **Phase 2 (ASHQ1 Engine)**: Dissects individual layer activations using multi-source importance matrices (`imatrix.gguf`, legacy `imatrix.dat` supported). Assigns precision tiers (`IQ2_XXS` through `Q8_0` and `F32`) dynamically based on layer sensitivity and tensor class, holding uncovered tensors at `IQ4_XS` or above.
### 2. Comprehensive Model Architecture Support
* **Dense & MoE Transformers**: Precise expert protection with token-routing stabilization.
* **Recurrent & Hybrid Models (GDN / Mamba / RWKV / Qwen3.5)**: Guaranteed `Q8_0` memory state retention to ensure long-context recurrence stability.
* **Multimodal Towers (CLIP / Vision Encoders)**: Dedicated `ASHQ1-mmproj.py` engine preserving spatial embeddings and layer-normalization vectors in `F32`/`F16`.
* **Multi-Token Prediction (MTP / NextN)**: Automatic extraction, isolation, and high-precision encoding (`Q6_K`/`Q8_0`) of speculative decoding heads, including nested projections (`nextn.eh_proj`, `nextn.embed_tokens`) and their `F32`-pinned norms.
---
## πŸ“Š Standardized ASHQ1 Tiers
All tiers maintain strict byte-budget percentages relative to the original unquantized BF16 model:
| Tier | File Ratio | Base Type | Typical Use Case | Target Preservation |
| :--- | :---: | :---: | :--- | :--- |
| **Nano** | **24%** | `IQ3_XXS` | Maximum compression, edge & mobile VRAM | Core gates `Q6_K`, Down-proj `IQ2_S` |
| **Mini** | **27%** | `IQ4_XS` | Efficient high-throughput serving | Balanced `IQ4_XS`/`IQ3_S` distribution |
| **Compact** | **33%** | `IQ4_XS` | Balanced daily-driver footprint | Down-proj `Q4_K`, Gate/Up `IQ4_XS` |
| **Quality** | **39%** | `Q5_K_M` | Near-lossless general deployment | Full `Q4_K`/`Q5_K` attention coverage |
| **Fidelity** | **48%** | `Q6_K` | Maximum analytical precision (raw BF16 lineage) | High-precision `Q5_K`/`Q6_K`/`Q8_0` mix |
*Note: Models originating from an AutoRound int4 lineage cap their weight allocations at `Q5_K`, as theoretical information saturation is fully realized. Attention gates settle at `Q6_K` and recurrent states at `Q8_0` on that lineage, and every tensor missing from the imatrix keeps `IQ4_XS` or above.*
> **Int4 lineage tier ladder**: `Compact` (33%) already drives every attention and FFN projection to the `Q5_K` cap. Because the perplexity gain beyond `Compact` is near-zero across all model sizes (1B to 9B, with Ξ” PPL ≀ 0.0358), `Compact` serves as the top tier on this lineage. Both `Quality` and `Fidelity` are skipped by default.
>
> Measured on a 9B `qwen35` source (17 091 MiB BF16): Nano **24.02%**, Mini **27.01%**, Compact **33.06%**.
### Perplexity Benchmarks (Ornith-1.5-9B)
Evaluated on `wiki.test.raw` (Wikitext-2), `n_ctx=2048`, 64 chunks, Flash-Attention enabled:
| Tier | Size | VRAM Budget | PPL | Ξ” vs Quality | Speed (RTX 8GB) |
|---|---|---|---|---|---|
| **Quality-36pc** | 6.06 GiB | ~7.5 GiB | **8.0932** | baseline | ~1241 tok/s |
| **Compact-33pc** | 5.65 GiB | ~7.0 GiB | **8.1290** | +0.0358 | ~1241 tok/s |
| **Mini-27pc** | 4.62 GiB | ~5.8 GiB | **9.5101** | +1.4169 | ~1442 tok/s |
| **Nano-24pc** | 4.01 GiB | ~4.5 GiB | **10.3148** | **+2.2216** | 1190.7 tok/s |
> **Run note:** The Nano-24pc result above is the 2026-08-20 validation run: 4,106 MiB on disk, 3.84 effective quantizer BPW, `ctx=2048`, 64 chunks, batch 512, 15 threads, and Flash-Attention.
>
> **Takeaways:**
> - `Quality-36pc` provides near-lossless perplexity for production inference.
> - `Compact-33pc` loses only **0.0358 PPL** while saving ~416 MiB, ideal for 8 GB VRAM setups.
> - `Mini-27pc` maintains strong conversational coherence under tight memory constraints.
### 🎯 Recommended Minimum Tiers by Model Size
Smaller parameter architectures require higher relative bit precision to prevent degradation of core reasoning representations:
* **β‰₯ 9B Parameters**: **Mini** (27% ratio) β€” Large parameter capacity preserves semantic integrity at lower bit rates.
* **~ 4B Parameters**: **Compact** (33% ratio) β€” Optimal balance between memory footprint and dense layer preservation.
* **~ 3B Parameters**: **Quality** (39% ratio) β€” Higher baseline precision protects critical routing and attention projections.
* **≀ 1B Parameters**: **Fidelity** (48% ratio) β€” Compact architectures require maximum parameter density.
---
## πŸ› οΈ Suite Components
| Script | Purpose |
| :--- | :--- |
| `00_SAFETENSORS-to-AutoRound-BF16-GGUF.py` | AutoRound tuner & streaming GGUF builder. Emits lineage provenance sidecars. |
| `00b_BF16-GGUF-MTP-extract.py` | Standalone speculative draft extractor for Multi-Token Prediction layers. |
| `01_create-calibration-dataset-and-imatrix.py` | End-to-end dataset builder (Agentic/Frontier/Logic) and GPU-autotuned `llama-imatrix` runner. |
| `01b_BF16-GGUF-modules-fusion.py` | Lossless merger combining base models, vision projectors (`mmproj`), and MTP heads. |
| `02_BF16-GGUF-to-ASHQ1.py` | Automated orchestrator executing batch quantization across all target tiers. |
| `03_perplexity_test.py` | Perplexity validation suite using `llama-perplexity` over reference corpora.
| `ASHQ1.py` | Core hybrid quantization optimizer with greedy knapsack utility scheduling and tied-weight detection. |
| `ASHQ1-mmproj.py` | Vision projector quantizer applying selective deep-block boosting and critical layer pinning. |
---
## ⚑ Quick Start
### 1. Requirements
Ensure CUDA, PyTorch, and `auto-round` are installed:
```bash
pip install auto-round torchvision safetensors gguf numpy huggingface_hub
```
### 2. End-to-End Workflow
```bash
# Step 0: Optimize safetensors and produce pristine AutoRound BF16 GGUF
python 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py ./safetensors/
# Step 1: Compute calibration activation statistics (imatrix)
python 01_create-calibration-dataset-and-imatrix.py
# Step 2: Generate all ASHQ1 standardized tiers
python 02_BF16-GGUF-to-ASHQ1.py
```
### 3. Recommended Inference Parameters
When serving ASHQ1 quantized models with `llama.cpp`, enable 4-bit KV cache quantization for optimal memory efficiency across extended context lengths:
```bash
llama-server -m model-AutoRound-ASHQ1-Quality-39pc.gguf -c 32768 --cache-type-k q4_0 --cache-type-v q4_0 -ngl 99
```
---
## πŸ“œ Citation & Credits
The AutoRound + ASHQ1 suite builds directly upon fundamental research and tooling across the open-source ecosystem:
* **ASHQ1 (Autonomous Selective Hybrid Quantization)** by **[wepiqx](https://huggingface.co/wepiqx/ASHQ1)**:
Original mathematical formulation of the priority-queue-driven knapsack optimizer, tied-group detection using numerical activation hashes, and theoretical MSE reduction scheduling.
* **Empero AI ([Qwen3.8-27B-Ridge](https://huggingface.co/empero-ai/Qwen3.8-27B-Ridge-GGUF))**:
Pioneering architectural insights on Gated-DeltaNet (GDN) hybrid attention preservation β€” specifically locking recurrence states (`ssm_alpha`, `ssm_beta`) in `Q8_0` and preserving native Multi-Token Prediction (MTP) draft heads.
* **Intel AutoRound**:
Sign-gradient-based optimization framework for low-bit weight reorganization with Hessian compensation.
* **llama.cpp** by **[Georgi Gerganov & ggml contributors](https://github.com/ggml-org/llama.cpp)**:
Core GGML/GGUF format definitions, runtime execution kernels, and quantization tools (`llama-quantize`, `llama-imatrix`).
* **Calibration Methodology & Recipes**:
Activation corpus curation inspired by **[Bartowski](https://huggingface.co/bartowski)** and multi-matrix combination techniques.