🧮 Paretrix GGUF quantizations of jinaai/ReaderLM-v2

Model MiB PPL ΔPPL KLD RMS Δp top-p Pareto
Q8_0 (stock) 1570 12.4136 +0.0284 0.0016 0.99% 97.9% ★
Fidelity-48pc 1422 12.4161 +0.0309 0.0031 1.39% 96.9% ★
Precision-42pc 1214 12.3712† −0.0140 0.0055 1.84% 95.9% ≡ Q6_K-imx
Q6_K-imx (stock) 1214 12.3712† −0.0140 0.0055 1.84% 95.9% ≡ Precision-42pc
Quality-36pc 1073 12.4244 +0.0392 0.0165 3.25% 93.1% ≡ Q5_K_M-imx
Q5_K_M-imx (stock) 1073 12.4244 +0.0392 0.0165 3.25% 93.1% ≡ Quality-36pc
Compact-33pc ⭐ 979 12.4617 +0.0765 0.0309 4.35% 90.9% ★
Mini-30pc 890 12.6359 +0.2507 0.0553 5.95% 87.7% ≈ IQ4_XS-imx (+0.0027 · −36 MiB)
IQ4_XS-imx (stock) 854 12.5393 +0.1541 0.0580 5.96% 87.6% ★
Nano-27pc 802 12.6418 +0.2566 0.0887 7.39% 84.9% ★
IQ3_M-imx (stock) 741 13.3774 +0.9922 0.1706 10.37% 79.6% ★

ℹ️ About Paretrix Quantization Suite

Paretrix is an empirical, activation-aware quantization suite for GGUF language models and speculative decoding modules.

Pareto + Matrix → Paretrix. Each tier is priced so that no better trade exists at that compression. The suite measures real activation sensitivity (ΔKLD/MiB) per tensor class through llama-imatrix, learns rate tables from cross-architecture campaigns (THE MATRIX), and allocates bitwidths under exact budget targets: flat recipes where uniformity wins, rate-calibrated knapsack where heterogeneity pays.

Standard uniform quantization applies one bitwidth family across the whole network. Paretrix reads the model's own activation field (which classes are expensive to cut, which are nearly free, and where depth matters), then spends the budget where measurements show the greatest return.

Downloads last month
298
GGUF
Model size
2B params
Architecture
qwen2
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Soulfate24/ReaderLM-v2-Paretrix

Quantized
(39)
this model