AutoRound + ASHQ1 Double-Quantization Suite
The AutoRound + ASHQ1 Suite delivers a complete pipeline for creating ultra-high-fidelity GGUF models. By combining gradient-guided weight reorganization (AutoRound W4A16) with fine-grained activation-aware tensor assignment (ASHQ1 Imatrix Engine), this suite establishes a new standard for low-bit LLM compression.
π Key Architecture & Highlights
βββββββββββββββββββββββββββ
β Safetensors (Raw / BF16)β
ββββββββββββββ¬βββββββββββββ
β 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py
βΌ
βββββββββββββββββββββββββββ
β AutoRound Optimization β βββΊ Iterative sign-rounding & Hessian estimation
ββββββββββββββ¬βββββββββββββ
β Streaming dequantization + GGUF encapsulation
βΌ
βββββββββββββββββββββββββββ
β AutoRound-Infused BF16 β βββΊ Lineage recorded in sidecar metadata
ββββββββββββββ¬βββββββββββββ
β 01_create-calibration-dataset-and-imatrix.py
βΌ
βββββββββββββββββββββββββββ
β Multi-Source Imatrix β βββΊ Agentic, Frontier, Logic & Diversity corpus
ββββββββββββββ¬βββββββββββββ
β 02_BF16-GGUF-to-ASHQ1.py (ASHQ1 Engine)
βΌ
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
β Standardized ASHQ1 Tiers: Nano β’ Mini β’ Compact β’ Quality β
βββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββββ
1. Dual-Phase Quantization Synergy
- Phase 1 (AutoRound W4A16): Reconditions original full-precision matrices via small-sample Hessian compensation. The optimized rounding directions remain intact when converted into BF16 GGUF containers.
- Phase 2 (ASHQ1 Engine): Dissects individual layer activations using multi-source importance matrices (
imatrix.gguf, legacyimatrix.datsupported). Assigns precision tiers (IQ2_XXSthroughQ8_0andF32) dynamically based on layer sensitivity and tensor class, holding uncovered tensors atIQ4_XSor above.
2. Comprehensive Model Architecture Support
- Dense & MoE Transformers: Precise expert protection with token-routing stabilization.
- Recurrent & Hybrid Models (GDN / Mamba / RWKV / Qwen3.5): Guaranteed
Q8_0memory state retention to ensure long-context recurrence stability. - Multimodal Towers (CLIP / Vision Encoders): Dedicated
ASHQ1-mmproj.pyengine preserving spatial embeddings and layer-normalization vectors inF32/F16. - Multi-Token Prediction (MTP / NextN): Automatic extraction, isolation, and high-precision encoding (
Q6_K/Q8_0) of speculative decoding heads, including nested projections (nextn.eh_proj,nextn.embed_tokens) and theirF32-pinned norms.
π Standardized ASHQ1 Tiers
All tiers maintain strict byte-budget percentages relative to the original unquantized BF16 model:
| Tier | File Ratio | Base Type | Typical Use Case | Target Preservation |
|---|---|---|---|---|
| Nano | 24% | IQ3_XXS |
Maximum compression, edge & mobile VRAM | Core gates Q6_K, Down-proj IQ2_S |
| Mini | 27% | IQ4_XS |
Efficient high-throughput serving | Balanced IQ4_XS/IQ3_S distribution |
| Compact | 33% | IQ4_XS |
Balanced daily-driver footprint | Down-proj Q4_K, Gate/Up IQ4_XS |
| Quality | 39% | Q5_K_M |
Near-lossless general deployment | Full Q4_K/Q5_K attention coverage |
| Fidelity | 48% | Q6_K |
Maximum analytical precision (raw BF16 lineage) | High-precision Q5_K/Q6_K/Q8_0 mix |
Note: Models originating from an AutoRound int4 lineage cap their weight allocations at Q5_K, as theoretical information saturation is fully realized. Attention gates settle at Q6_K and recurrent states at Q8_0 on that lineage, and every tensor missing from the imatrix keeps IQ4_XS or above.
Int4 lineage tier ladder:
Compact(33%) already drives every attention and FFN projection to theQ5_Kcap. Because the perplexity gain beyondCompactis near-zero across all model sizes (1B to 9B, with Ξ PPL β€ 0.0358),Compactserves as the top tier on this lineage. BothQualityandFidelityare skipped by default.Measured on a 9B
qwen35source (17 091 MiB BF16): Nano 24.02%, Mini 27.01%, Compact 33.06%.
Perplexity Benchmarks (Ornith-1.5-9B)
Evaluated on wiki.test.raw (Wikitext-2), n_ctx=2048, 64 chunks, Flash-Attention enabled:
| Tier | Size | VRAM Budget | PPL | Ξ vs Quality | Speed (RTX 8GB) |
|---|---|---|---|---|---|
| Quality-36pc | 6.06 GiB | ~7.5 GiB | 8.0932 | baseline | ~1241 tok/s |
| Compact-33pc | 5.65 GiB | ~7.0 GiB | 8.1290 | +0.0358 | ~1241 tok/s |
| Mini-27pc | 4.62 GiB | ~5.8 GiB | 9.5101 | +1.4169 | ~1442 tok/s |
| Nano-24pc | 4.01 GiB | ~4.5 GiB | 10.3148 | +2.2216 | 1190.7 tok/s |
Run note: The Nano-24pc result above is the 2026-08-20 validation run: 4,106 MiB on disk, 3.84 effective quantizer BPW,
ctx=2048, 64 chunks, batch 512, 15 threads, and Flash-Attention.Takeaways:
Quality-36pcprovides near-lossless perplexity for production inference.Compact-33pcloses only 0.0358 PPL while saving ~416 MiB, ideal for 8 GB VRAM setups.Mini-27pcmaintains strong conversational coherence under tight memory constraints.
π― Recommended Minimum Tiers by Model Size
Smaller parameter architectures require higher relative bit precision to prevent degradation of core reasoning representations:
- β₯ 9B Parameters: Mini (27% ratio) β Large parameter capacity preserves semantic integrity at lower bit rates.
- ~ 4B Parameters: Compact (33% ratio) β Optimal balance between memory footprint and dense layer preservation.
- ~ 3B Parameters: Quality (39% ratio) β Higher baseline precision protects critical routing and attention projections.
- β€ 1B Parameters: Fidelity (48% ratio) β Compact architectures require maximum parameter density.
π οΈ Suite Components
| Script | Purpose |
|---|---|
00_SAFETENSORS-to-AutoRound-BF16-GGUF.py |
AutoRound tuner & streaming GGUF builder. Emits lineage provenance sidecars. |
00b_BF16-GGUF-MTP-extract.py |
Standalone speculative draft extractor for Multi-Token Prediction layers. |
01_create-calibration-dataset-and-imatrix.py |
End-to-end dataset builder (Agentic/Frontier/Logic) and GPU-autotuned llama-imatrix runner. |
01b_BF16-GGUF-modules-fusion.py |
Lossless merger combining base models, vision projectors (mmproj), and MTP heads. |
02_BF16-GGUF-to-ASHQ1.py |
Automated orchestrator executing batch quantization across all target tiers. |
03_perplexity_test.py |
Perplexity validation suite using llama-perplexity over reference corpora. |
ASHQ1.py |
Core hybrid quantization optimizer with greedy knapsack utility scheduling and tied-weight detection. |
ASHQ1-mmproj.py |
Vision projector quantizer applying selective deep-block boosting and critical layer pinning. |
β‘ Quick Start
1. Requirements
Ensure CUDA, PyTorch, and auto-round are installed:
pip install auto-round torchvision safetensors gguf numpy huggingface_hub
2. End-to-End Workflow
# Step 0: Optimize safetensors and produce pristine AutoRound BF16 GGUF
python 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py ./safetensors/
# Step 1: Compute calibration activation statistics (imatrix)
python 01_create-calibration-dataset-and-imatrix.py
# Step 2: Generate all ASHQ1 standardized tiers
python 02_BF16-GGUF-to-ASHQ1.py
3. Recommended Inference Parameters
When serving ASHQ1 quantized models with llama.cpp, enable 4-bit KV cache quantization for optimal memory efficiency across extended context lengths:
llama-server -m model-AutoRound-ASHQ1-Quality-39pc.gguf -c 32768 --cache-type-k q4_0 --cache-type-v q4_0 -ngl 99
π Citation & Credits
The AutoRound + ASHQ1 suite builds directly upon fundamental research and tooling across the open-source ecosystem:
- ASHQ1 (Autonomous Selective Hybrid Quantization) by wepiqx: Original mathematical formulation of the priority-queue-driven knapsack optimizer, tied-group detection using numerical activation hashes, and theoretical MSE reduction scheduling.
- Empero AI (Qwen3.8-27B-Ridge):
Pioneering architectural insights on Gated-DeltaNet (GDN) hybrid attention preservation β specifically locking recurrence states (
ssm_alpha,ssm_beta) inQ8_0and preserving native Multi-Token Prediction (MTP) draft heads. - Intel AutoRound: Sign-gradient-based optimization framework for low-bit weight reorganization with Hessian compensation.
- llama.cpp by Georgi Gerganov & ggml contributors:
Core GGML/GGUF format definitions, runtime execution kernels, and quantization tools (
llama-quantize,llama-imatrix). - Calibration Methodology & Recipes: Activation corpus curation inspired by Bartowski and multi-matrix combination techniques.