AutoRound + ASHQ1 Double-Quantization Suite

The AutoRound + ASHQ1 Suite delivers a complete pipeline for creating ultra-high-fidelity GGUF models. By combining gradient-guided weight reorganization (AutoRound W4A16) with fine-grained activation-aware tensor assignment (ASHQ1 Imatrix Engine), this suite establishes a new standard for low-bit LLM compression.


🌟 Key Architecture & Highlights

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Safetensors (Raw / BF16)β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚ 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ AutoRound Optimization  β”‚ ──► Iterative sign-rounding & Hessian estimation
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚ Streaming dequantization + GGUF encapsulation
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ AutoRound-Infused BF16  β”‚ ──► Lineage recorded in sidecar metadata
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚ 01_create-calibration-dataset-and-imatrix.py
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Multi-Source Imatrix    β”‚ ──► Agentic, Frontier, Logic & Diversity corpus
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
             β”‚ 02_BF16-GGUF-to-ASHQ1.py (ASHQ1 Engine)
             β–Ό
β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚ Standardized ASHQ1 Tiers: Nano β€’ Mini β€’ Compact β€’ Quality   β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜

1. Dual-Phase Quantization Synergy

  • Phase 1 (AutoRound W4A16): Reconditions original full-precision matrices via small-sample Hessian compensation. The optimized rounding directions remain intact when converted into BF16 GGUF containers.
  • Phase 2 (ASHQ1 Engine): Dissects individual layer activations using multi-source importance matrices (imatrix.gguf, legacy imatrix.dat supported). Assigns precision tiers (IQ2_XXS through Q8_0 and F32) dynamically based on layer sensitivity and tensor class, holding uncovered tensors at IQ4_XS or above.

2. Comprehensive Model Architecture Support

  • Dense & MoE Transformers: Precise expert protection with token-routing stabilization.
  • Recurrent & Hybrid Models (GDN / Mamba / RWKV / Qwen3.5): Guaranteed Q8_0 memory state retention to ensure long-context recurrence stability.
  • Multimodal Towers (CLIP / Vision Encoders): Dedicated ASHQ1-mmproj.py engine preserving spatial embeddings and layer-normalization vectors in F32/F16.
  • Multi-Token Prediction (MTP / NextN): Automatic extraction, isolation, and high-precision encoding (Q6_K/Q8_0) of speculative decoding heads, including nested projections (nextn.eh_proj, nextn.embed_tokens) and their F32-pinned norms.

πŸ“Š Standardized ASHQ1 Tiers

All tiers maintain strict byte-budget percentages relative to the original unquantized BF16 model:

Tier File Ratio Base Type Typical Use Case Target Preservation
Nano 24% IQ3_XXS Maximum compression, edge & mobile VRAM Core gates Q6_K, Down-proj IQ2_S
Mini 27% IQ4_XS Efficient high-throughput serving Balanced IQ4_XS/IQ3_S distribution
Compact 33% IQ4_XS Balanced daily-driver footprint Down-proj Q4_K, Gate/Up IQ4_XS
Quality 39% Q5_K_M Near-lossless general deployment Full Q4_K/Q5_K attention coverage
Fidelity 48% Q6_K Maximum analytical precision (raw BF16 lineage) High-precision Q5_K/Q6_K/Q8_0 mix

Note: Models originating from an AutoRound int4 lineage cap their weight allocations at Q5_K, as theoretical information saturation is fully realized. Attention gates settle at Q6_K and recurrent states at Q8_0 on that lineage, and every tensor missing from the imatrix keeps IQ4_XS or above.

Int4 lineage tier ladder: Compact (33%) already drives every attention and FFN projection to the Q5_K cap. Because the perplexity gain beyond Compact is near-zero across all model sizes (1B to 9B, with Ξ” PPL ≀ 0.0358), Compact serves as the top tier on this lineage. Both Quality and Fidelity are skipped by default.

Measured on a 9B qwen35 source (17 091 MiB BF16): Nano 24.02%, Mini 27.01%, Compact 33.06%.

Perplexity Benchmarks (Ornith-1.5-9B)

Evaluated on wiki.test.raw (Wikitext-2), n_ctx=2048, 64 chunks, Flash-Attention enabled:

Tier Size VRAM Budget PPL Ξ” vs Quality Speed (RTX 8GB)
Quality-36pc 6.06 GiB ~7.5 GiB 8.0932 baseline ~1241 tok/s
Compact-33pc 5.65 GiB ~7.0 GiB 8.1290 +0.0358 ~1241 tok/s
Mini-27pc 4.62 GiB ~5.8 GiB 9.5101 +1.4169 ~1442 tok/s
Nano-24pc 4.01 GiB ~4.5 GiB 10.3148 +2.2216 1190.7 tok/s

Run note: The Nano-24pc result above is the 2026-08-20 validation run: 4,106 MiB on disk, 3.84 effective quantizer BPW, ctx=2048, 64 chunks, batch 512, 15 threads, and Flash-Attention.

Takeaways:

  • Quality-36pc provides near-lossless perplexity for production inference.
  • Compact-33pc loses only 0.0358 PPL while saving ~416 MiB, ideal for 8 GB VRAM setups.
  • Mini-27pc maintains strong conversational coherence under tight memory constraints.

🎯 Recommended Minimum Tiers by Model Size

Smaller parameter architectures require higher relative bit precision to prevent degradation of core reasoning representations:

  • β‰₯ 9B Parameters: Mini (27% ratio) β€” Large parameter capacity preserves semantic integrity at lower bit rates.
  • ~ 4B Parameters: Compact (33% ratio) β€” Optimal balance between memory footprint and dense layer preservation.
  • ~ 3B Parameters: Quality (39% ratio) β€” Higher baseline precision protects critical routing and attention projections.
  • ≀ 1B Parameters: Fidelity (48% ratio) β€” Compact architectures require maximum parameter density.

πŸ› οΈ Suite Components

Script Purpose
00_SAFETENSORS-to-AutoRound-BF16-GGUF.py AutoRound tuner & streaming GGUF builder. Emits lineage provenance sidecars.
00b_BF16-GGUF-MTP-extract.py Standalone speculative draft extractor for Multi-Token Prediction layers.
01_create-calibration-dataset-and-imatrix.py End-to-end dataset builder (Agentic/Frontier/Logic) and GPU-autotuned llama-imatrix runner.
01b_BF16-GGUF-modules-fusion.py Lossless merger combining base models, vision projectors (mmproj), and MTP heads.
02_BF16-GGUF-to-ASHQ1.py Automated orchestrator executing batch quantization across all target tiers.
03_perplexity_test.py Perplexity validation suite using llama-perplexity over reference corpora.
ASHQ1.py Core hybrid quantization optimizer with greedy knapsack utility scheduling and tied-weight detection.
ASHQ1-mmproj.py Vision projector quantizer applying selective deep-block boosting and critical layer pinning.

⚑ Quick Start

1. Requirements

Ensure CUDA, PyTorch, and auto-round are installed:

pip install auto-round torchvision safetensors gguf numpy huggingface_hub

2. End-to-End Workflow

# Step 0: Optimize safetensors and produce pristine AutoRound BF16 GGUF
python 00_SAFETENSORS-to-AutoRound-BF16-GGUF.py ./safetensors/

# Step 1: Compute calibration activation statistics (imatrix)
python 01_create-calibration-dataset-and-imatrix.py

# Step 2: Generate all ASHQ1 standardized tiers
python 02_BF16-GGUF-to-ASHQ1.py

3. Recommended Inference Parameters

When serving ASHQ1 quantized models with llama.cpp, enable 4-bit KV cache quantization for optimal memory efficiency across extended context lengths:

llama-server -m model-AutoRound-ASHQ1-Quality-39pc.gguf -c 32768 --cache-type-k q4_0 --cache-type-v q4_0 -ngl 99

πŸ“œ Citation & Credits

The AutoRound + ASHQ1 suite builds directly upon fundamental research and tooling across the open-source ecosystem:

  • ASHQ1 (Autonomous Selective Hybrid Quantization) by wepiqx: Original mathematical formulation of the priority-queue-driven knapsack optimizer, tied-group detection using numerical activation hashes, and theoretical MSE reduction scheduling.
  • Empero AI (Qwen3.8-27B-Ridge): Pioneering architectural insights on Gated-DeltaNet (GDN) hybrid attention preservation β€” specifically locking recurrence states (ssm_alpha, ssm_beta) in Q8_0 and preserving native Multi-Token Prediction (MTP) draft heads.
  • Intel AutoRound: Sign-gradient-based optimization framework for low-bit weight reorganization with Hessian compensation.
  • llama.cpp by Georgi Gerganov & ggml contributors: Core GGML/GGUF format definitions, runtime execution kernels, and quantization tools (llama-quantize, llama-imatrix).
  • Calibration Methodology & Recipes: Activation corpus curation inspired by Bartowski and multi-matrix combination techniques.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support