--- license: apache-2.0 tags: - quantization - gguf - autoround - imatrix - hybrid-quantization - llama.cpp - text-generation --- # AutoRound + ASHQ1 Double-Quantization Suite The **AutoRound + ASHQ1 Suite** delivers a complete pipeline for creating ultra-high-fidelity GGUF models. By combining gradient-guided weight reorganization (**AutoRound W4A16**) with fine-grained activation-aware tensor assignment (**ASHQ1 Imatrix Engine**), this suite establishes a new standard for low-bit LLM compression. --- ## 🌟 Key Architecture & Highlights ``` ┌─────────────────────────┐ │ Safetensors (Raw / BF16)│ └────────────┬────────────┘ │ 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py ▼ ┌─────────────────────────┐ │ AutoRound Optimization │ ──► Iterative sign-rounding & Hessian estimation └────────────┬────────────┘ │ Streaming dequantization + GGUF encapsulation ▼ ┌─────────────────────────┐ │ AutoRound-Infused BF16 │ ──► Lineage recorded in sidecar metadata └────────────┬────────────┘ │ 01_create-calibration-dataset-and-imatrix.py ▼ ┌─────────────────────────┐ │ Multi-Source Imatrix │ ──► Agentic, Frontier, Logic & Diversity corpus └────────────┬────────────┘ │ 02_BF16-GGUF-to-ASHQ1.py (ASHQ1 Engine) ▼ ┌─────────────────────────────────────────────────────────────┐ │ Standardized ASHQ1 Tiers: Nano • Mini • Compact • Quality │ └─────────────────────────────────────────────────────────────┘ ``` ### 1. Dual-Phase Quantization Synergy * **Phase 1 (AutoRound W4A16)**: Reconditions original full-precision matrices via small-sample Hessian compensation. The optimized rounding directions remain intact when converted into BF16 GGUF containers. * **Phase 2 (ASHQ1 Engine)**: Dissects individual layer activations using multi-source importance matrices (`imatrix.gguf`, legacy `imatrix.dat` supported). Assigns precision tiers (`IQ2_XXS` through `Q8_0` and `F32`) dynamically based on layer sensitivity and tensor class, holding uncovered tensors at `IQ4_XS` or above. ### 2. Comprehensive Model Architecture Support * **Dense & MoE Transformers**: Precise expert protection with token-routing stabilization. * **Recurrent & Hybrid Models (GDN / Mamba / RWKV / Qwen3.5)**: Guaranteed `Q8_0` memory state retention to ensure long-context recurrence stability. * **Multimodal Towers (CLIP / Vision Encoders)**: Dedicated `ASHQ1-mmproj.py` engine preserving spatial embeddings and layer-normalization vectors in `F32`/`F16`. * **Multi-Token Prediction (MTP / NextN)**: Automatic extraction, isolation, and high-precision encoding (`Q6_K`/`Q8_0`) of speculative decoding heads, including nested projections (`nextn.eh_proj`, `nextn.embed_tokens`) and their `F32`-pinned norms. --- ## 📊 Standardized ASHQ1 Tiers All tiers maintain strict byte-budget percentages relative to the original unquantized BF16 model: | Tier | File Ratio | Base Type | Typical Use Case | Target Preservation | | :--- | :---: | :---: | :--- | :--- | | **Nano** | **24%** | `IQ3_XXS` | Maximum compression, edge & mobile VRAM | Core gates `Q6_K`, Down-proj `IQ2_S` | | **Mini** | **27%** | `IQ4_XS` | Efficient high-throughput serving | Balanced `IQ4_XS`/`IQ3_S` distribution | | **Compact** | **33%** | `IQ4_XS` | Balanced daily-driver footprint | Down-proj `Q4_K`, Gate/Up `IQ4_XS` | | **Quality** | **39%** | `Q5_K_M` | Near-lossless general deployment | Full `Q4_K`/`Q5_K` attention coverage | | **Fidelity** | **48%** | `Q6_K` | Maximum analytical precision (raw BF16 lineage) | High-precision `Q5_K`/`Q6_K`/`Q8_0` mix | *Note: Models originating from an AutoRound int4 lineage cap their weight allocations at `Q5_K`, as theoretical information saturation is fully realized. Attention gates settle at `Q6_K` and recurrent states at `Q8_0` on that lineage, and every tensor missing from the imatrix keeps `IQ4_XS` or above.* > **Int4 lineage tier ladder**: `Compact` (33%) already drives every attention and FFN projection to the `Q5_K` cap. Because the perplexity gain beyond `Compact` is near-zero across all model sizes (1B to 9B, with Δ PPL ≤ 0.0358), `Compact` serves as the top tier on this lineage. Both `Quality` and `Fidelity` are skipped by default. > > Measured on a 9B `qwen35` source (17 091 MiB BF16): Nano **24.02%**, Mini **27.01%**, Compact **33.06%**. ### Perplexity Benchmarks (Ornith-1.5-9B) Evaluated on `wiki.test.raw` (Wikitext-2), `n_ctx=2048`, 64 chunks, Flash-Attention enabled: | Tier | Size | VRAM Budget | PPL | Δ vs Quality | Speed (RTX 8GB) | |---|---|---|---|---|---| | **Quality-36pc** | 6.06 GiB | ~7.5 GiB | **8.0932** | baseline | ~1241 tok/s | | **Compact-33pc** | 5.65 GiB | ~7.0 GiB | **8.1290** | +0.0358 | ~1241 tok/s | | **Mini-27pc** | 4.62 GiB | ~5.8 GiB | **9.5101** | +1.4169 | ~1442 tok/s | | **Nano-24pc** | 4.01 GiB | ~4.5 GiB | **10.3148** | **+2.2216** | 1190.7 tok/s | > **Run note:** The Nano-24pc result above is the 2026-08-20 validation run: 4,106 MiB on disk, 3.84 effective quantizer BPW, `ctx=2048`, 64 chunks, batch 512, 15 threads, and Flash-Attention. > > **Takeaways:** > - `Quality-36pc` provides near-lossless perplexity for production inference. > - `Compact-33pc` loses only **0.0358 PPL** while saving ~416 MiB, ideal for 8 GB VRAM setups. > - `Mini-27pc` maintains strong conversational coherence under tight memory constraints. ### 🎯 Recommended Minimum Tiers by Model Size Smaller parameter architectures require higher relative bit precision to prevent degradation of core reasoning representations: * **≥ 9B Parameters**: **Mini** (27% ratio) — Large parameter capacity preserves semantic integrity at lower bit rates. * **~ 4B Parameters**: **Compact** (33% ratio) — Optimal balance between memory footprint and dense layer preservation. * **~ 3B Parameters**: **Quality** (39% ratio) — Higher baseline precision protects critical routing and attention projections. * **≤ 1B Parameters**: **Fidelity** (48% ratio) — Compact architectures require maximum parameter density. --- ## 🛠️ Suite Components | Script | Purpose | | :--- | :--- | | `00_SAFETENSORS-to-AutoRound-BF16-GGUF.py` | AutoRound tuner & streaming GGUF builder. Emits lineage provenance sidecars. | | `00b_BF16-GGUF-MTP-extract.py` | Standalone speculative draft extractor for Multi-Token Prediction layers. | | `01_create-calibration-dataset-and-imatrix.py` | End-to-end dataset builder (Agentic/Frontier/Logic) and GPU-autotuned `llama-imatrix` runner. | | `01b_BF16-GGUF-modules-fusion.py` | Lossless merger combining base models, vision projectors (`mmproj`), and MTP heads. | | `02_BF16-GGUF-to-ASHQ1.py` | Automated orchestrator executing batch quantization across all target tiers. | | `03_perplexity_test.py` | Perplexity validation suite using `llama-perplexity` over reference corpora. | `ASHQ1.py` | Core hybrid quantization optimizer with greedy knapsack utility scheduling and tied-weight detection. | | `ASHQ1-mmproj.py` | Vision projector quantizer applying selective deep-block boosting and critical layer pinning. | --- ## ⚡ Quick Start ### 1. Requirements Ensure CUDA, PyTorch, and `auto-round` are installed: ```bash pip install auto-round torchvision safetensors gguf numpy huggingface_hub ``` ### 2. End-to-End Workflow ```bash # Step 0: Optimize safetensors and produce pristine AutoRound BF16 GGUF python 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py ./safetensors/ # Step 1: Compute calibration activation statistics (imatrix) python 01_create-calibration-dataset-and-imatrix.py # Step 2: Generate all ASHQ1 standardized tiers python 02_BF16-GGUF-to-ASHQ1.py ``` ### 3. Recommended Inference Parameters When serving ASHQ1 quantized models with `llama.cpp`, enable 4-bit KV cache quantization for optimal memory efficiency across extended context lengths: ```bash llama-server -m model-AutoRound-ASHQ1-Quality-39pc.gguf -c 32768 --cache-type-k q4_0 --cache-type-v q4_0 -ngl 99 ``` --- ## 📜 Citation & Credits The AutoRound + ASHQ1 suite builds directly upon fundamental research and tooling across the open-source ecosystem: * **ASHQ1 (Autonomous Selective Hybrid Quantization)** by **[wepiqx](https://huggingface.co/wepiqx/ASHQ1)**: Original mathematical formulation of the priority-queue-driven knapsack optimizer, tied-group detection using numerical activation hashes, and theoretical MSE reduction scheduling. * **Empero AI ([Qwen3.8-27B-Ridge](https://huggingface.co/empero-ai/Qwen3.8-27B-Ridge-GGUF))**: Pioneering architectural insights on Gated-DeltaNet (GDN) hybrid attention preservation — specifically locking recurrence states (`ssm_alpha`, `ssm_beta`) in `Q8_0` and preserving native Multi-Token Prediction (MTP) draft heads. * **Intel AutoRound**: Sign-gradient-based optimization framework for low-bit weight reorganization with Hessian compensation. * **llama.cpp** by **[Georgi Gerganov & ggml contributors](https://github.com/ggml-org/llama.cpp)**: Core GGML/GGUF format definitions, runtime execution kernels, and quantization tools (`llama-quantize`, `llama-imatrix`). * **Calibration Methodology & Recipes**: Activation corpus curation inspired by **[Bartowski](https://huggingface.co/bartowski)** and multi-matrix combination techniques.