CompressedGemma commited on
Commit
f811774
Β·
verified Β·
1 Parent(s): 3ba66cb

Update README.md

Browse files
Files changed (1) hide show
  1. README.md +452 -463
README.md CHANGED
@@ -2,689 +2,696 @@
2
  license: mit
3
  ---
4
 
5
- # Disclaimer
6
-
7
- With recent changes, the runtime of a quantization operation has gone through the roof, but models quantized this way benefit from a massive increase in model coherency.
8
-
9
- I am not going to revert these changes, but will attempt to streamline this further.
10
-
11
- To use this properly you must dump a .hxh file and enable --hessian and --analog-imatrix, otherwise it will fall back to the typical RMSE path which is still vastly superior to everything else.
12
-
13
- It is worth the cost of using --hessian and --analog-imatrix, though.
14
-
15
-
16
  # HPC-Quantize
17
 
18
- ### Holographic Phase Contraction for Ultra-Low-Bit LLM Quantization
19
 
20
- **HPC-Quantize** is an MIT-licensed research quantization engine for aggressively compressing large language models while preserving inference behavior at extremely low bitrates.
21
 
22
- HPC was originally developed around a quantum-inspired Shor/Griffiths–Niu measurement formulation. The current engine has evolved beyond that approach: the production quantization path uses a **sequential Sieve**, bounded state back-action, and a **36-state Viterbi optimizer** for Q2 quantization.
23
 
24
- The central idea remains the same:
25
 
26
- > **Quantization error is not merely a collection of independent scalar errors. Its structure matters.**
27
 
28
- Rather than optimizing every block independently, HPC treats quantization as a structured discrete-state optimization problem.
29
 
30
  ---
31
 
32
- # What makes HPC different?
33
 
34
- Conventional low-bit quantization generally asks:
35
 
36
- $$
37
- \hat W_i =
38
- \arg\min_{\hat W}
39
- \|W_i-\hat W\|^2
40
- $$
41
-
42
- for each block independently.
43
 
44
- HPC instead introduces an explicit state space for quantization candidates and considers:
45
 
46
- * candidate reconstruction error
47
- * anisotropic error structure
48
- * neighboring quantization states
49
- * state diversity
50
- * candidate probability
51
- * cross-block state transitions
52
- * global sequence coherence
53
 
54
- The resulting optimization is closer to a graphical-model / sequence-decoding problem than conventional round-to-nearest quantization.
 
 
 
 
 
 
 
 
55
 
56
- Conceptually:
57
 
58
  ```text
59
- Original weights
60
- β”‚
61
- β–Ό
62
- Candidate generation
63
- β”‚
64
- β–Ό
65
- Error-aware scoring
66
- β”‚
67
- β–Ό
68
- Six-state HPC representation
69
- β”‚
70
- β–Ό
71
- Sequential Sieve
72
- β”‚
73
- β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”΄β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
74
- β”‚ β”‚
75
- state selection neighbor back-action
76
- β”‚ β”‚
77
- β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
78
- β–Ό
79
- Q2 state lattice
80
- β”‚
81
- β–Ό
82
- 36-state Viterbi
83
- β”‚
84
- β–Ό
85
- Candidate reconstruction
86
- β”‚
87
- β–Ό
88
- GGUF output
89
  ```
90
 
 
 
91
  ---
92
 
93
- # The HPC state space
94
 
95
- HPC represents quantization candidates using a six-state symbolic space:
96
 
97
- $$
98
- d\in\{0,1,2,3,4,5\}.
99
- $$
100
 
101
- The six states form the basic HPC **quhit**.
102
 
103
- Candidate parameters are mapped into these six symbolic states through the candidate-to-quhit mapping.
104
 
105
- This provides a compact representation of the candidate landscape while retaining multiple competing possibilities rather than immediately selecting the local minimum.
 
 
 
 
 
 
 
 
 
106
 
107
  ---
108
 
109
- # Q2: the 36-state space
110
 
111
- Q2 quantization has two coupled parameters:
112
 
113
- $$
114
- (d,d_{\min})
115
- $$
 
 
 
 
 
 
116
 
117
- Each parameter is projected into six HPC states.
118
 
119
- Consequently, the joint state is
120
 
121
- $$
122
- (q_D,q_M)
123
- $$
124
 
125
- with
126
 
127
  $$
128
- q_D,q_M\in\{0,\ldots,5\}.
129
  $$
130
 
131
- The complete symbolic Q2 state space is therefore:
 
 
132
 
133
  $$
134
- 6\times6=36.
 
135
  $$
136
 
137
- The state index is:
138
 
139
- $$
140
- s=6q_D+q_M.
141
- $$
142
 
143
- Thus:
144
 
145
- ```text
146
- dmin
147
- 0 1 2 3 4 5
148
- β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
149
- 0 β”‚ 0 1 2 3 4 5
150
- 1 β”‚ 6 7 8 9 10 11
151
- d 2 β”‚12 13 14 15 16 17
152
- 3 β”‚18 19 20 21 22 23
153
- 4 β”‚24 25 26 27 28 29
154
- 5 β”‚30 31 32 33 34 35
155
- ```
156
 
157
- Each symbolic state corresponds to the best physical Q2 candidate belonging to that state.
158
 
159
  ---
160
 
161
- # Error β†’ probability
162
 
163
- HPC converts candidate error into a Boltzmann-like amplitude:
164
 
165
- $$
166
- A_i =
167
- \exp\left(
168
- -\frac{E_i-E_{\min}}{2T}
169
- \right).
170
- $$
171
 
172
- Squaring gives the corresponding probability:
173
 
174
  $$
175
- P_i\propto
176
- \exp\left(
177
- -\frac{E_i-E_{\min}}{T}
178
- \right).
179
  $$
180
 
181
- This means the quantizer retains information about the *relative quality* of candidates rather than immediately discarding everything except the MSE winner.
182
 
183
- ---
184
 
185
- # The Sieve
186
 
187
- The current HPC engine replaces the original Shor/Griffiths–Niu measurement path with a classical sequential Sieve.
 
 
 
 
 
 
 
 
 
 
 
188
 
189
- For a state \(d\), its initial probability is:
190
 
191
- $$
192
- p_0(d)=|\alpha(d)|^2.
193
- $$
194
 
195
- Each neighboring state contributes a compatibility factor.
196
 
197
- For neighbor \(j\):
198
 
199
  $$
200
- B_j(d)
201
- =
202
- \sum_w p_j(w)
203
- \begin{cases}
204
- 0.85,&w=d\\
205
- 1,&w\ne d.
206
- \end{cases}
207
  $$
208
 
209
- Because the probabilities sum to one:
 
 
210
 
211
  $$
212
- B_j(d)=1-0.15p_j(d).
213
  $$
214
 
215
- The resulting score is therefore:
 
 
216
 
217
  $$
218
- S(d)=
219
- p_0(d)
220
- \prod_{j\in N}
221
- \left(1-0.15p_j(d)\right).
222
  $$
223
 
224
- The implementation evaluates this in logarithmic form:
225
 
226
  $$
227
- \ell(d)=
228
- \log p_0(d)
229
- +
230
- \sum_{j\in N}
231
- \log\left(1-0.15p_j(d)\right).
232
  $$
233
 
234
- This gives the Sieve a simple interpretation:
235
 
236
- > **A locally strong state remains strong, but states already strongly represented by neighboring blocks are gently penalized.**
237
-
238
- ---
239
-
240
- # Why 0.85?
241
-
242
- The factor:
 
 
 
 
243
 
244
- $$
245
- 0.85
246
- $$
247
 
248
- controls the strength of the local anti-correlation.
249
 
250
- If a neighboring site is completely concentrated on state \(d\),
251
 
252
- $$
253
- p_j(d)=1,
254
- $$
255
 
256
- then:
257
 
258
  $$
259
- B_j(d)=0.85.
260
  $$
261
 
262
- In log space this is:
263
 
264
- $$
265
- \log(0.85)\approx-0.1625.
266
- $$
267
 
268
- The Sieve therefore discourages repeated neighboring states without making them impossible.
269
 
270
- This is intentional.
271
-
272
- HPC does **not** attempt to force an artificial alternating pattern. It merely introduces a preference for diverse state assignments when alternative candidates remain competitive.
273
 
274
  ---
275
 
276
- # The 1.5 Sieve slack
277
 
278
- After calculating the six state scores, HPC finds:
279
 
280
- $$
281
- \ell_{\max}=\max_d\ell(d).
282
- $$
283
 
284
- A state survives if:
285
 
286
  $$
287
- \ell(d)\geq\ell_{\max}-1.5.
288
  $$
289
 
290
- Equivalently:
291
 
292
  $$
293
- \frac{S(d)}{S_{\max}}
294
- \geq e^{-1.5}
295
- \approx0.2231.
296
  $$
297
 
298
- Thus alternatives whose score is at least approximately **22.3% of the best candidate** remain eligible.
299
 
300
- This prevents premature collapse of the candidate landscape.
301
 
302
- The Sieve therefore acts as a **soft constraint**, not a hard nearest-neighbor rule.
303
 
304
  ---
305
 
306
- # Sequential collapse and back-action
307
 
308
- Once the current state is selected:
309
 
310
- $$
311
- d_k=\arg\max_d S_k(d),
312
- $$
313
 
314
- that state is collapsed.
315
 
316
- The selected state is then propagated into its remaining neighbors.
 
 
 
 
 
 
 
 
317
 
318
- For every live neighbor:
319
 
320
- $$
321
- \alpha_j(d_k)
322
- \leftarrow
323
- 0.85\alpha_j(d_k).
324
- $$
325
 
326
- The neighbor is then renormalized.
327
 
328
- This creates sequential conditioning:
329
 
330
  $$
331
  P(d_1)
332
  \rightarrow
333
- P(d_2\mid d_1)
334
  \rightarrow
335
- P(d_3\mid d_1,d_2)
336
- \rightarrow\cdots
337
  $$
338
 
339
- The graph therefore does not make independent decisions.
340
-
341
- Earlier decisions influence later decisions.
342
 
343
  ---
344
 
345
- # Q2 Viterbi optimization
346
 
347
- After the Sieve produces the six-state distributions, Q2 reconstructs the full 36-state lattice.
348
 
349
- For each symbolic state:
350
 
351
- $$
352
- s=(q_D,q_M),
353
- $$
354
-
355
- HPC retains the lowest-error physical candidate associated with that state.
356
-
357
- The Sieve distributions provide a prior:
358
 
359
  $$
360
- P(s)
361
- \approx
362
- P_D(q_D)P_M(q_M).
363
  $$
364
 
365
- The local Viterbi cost is:
366
 
367
- $$
368
- C_i(s)=
369
- E_i(s)
370
- -
371
- 0.25\bar E\log P_i(s).
372
- $$
373
-
374
- HPC then imposes a transition cost between adjacent blocks:
375
 
376
  $$
377
- T(s',s)=
378
- 0.08\bar E
379
- \left(
380
- |q_D-q_D'|
381
- +
382
- |q_M-q_M'|
383
- \right).
384
  $$
385
 
386
- This is Manhattan distance on the 6Γ—6 Q2 lattice.
387
 
388
- The dynamic program is:
389
 
390
  $$
391
- DP_i(s)=
392
- C_i(s)+
 
 
393
  \min_{s'}
394
  \left[
395
  DP_{i-1}(s')+T(s',s)
396
  \right].
397
  $$
398
 
399
- The result is a globally optimized sequence of Q2 states rather than independently optimized blocks.
 
 
400
 
401
  ---
402
 
403
- # Greedy safety override
404
 
405
- HPC does not blindly trust the global optimizer.
406
 
407
- After Viterbi selection, the engine compares the selected candidate against the best local reconstruction candidate.
408
 
409
- If the greedy candidate is sufficiently better:
410
 
411
- $$
412
- E_{\text{greedy}}
413
- <
414
- 0.95E_{\text{selected}},
415
- $$
416
 
417
- the local candidate wins.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
418
 
419
- This prevents global regularization from sacrificing too much actual reconstruction quality.
420
 
421
  ---
422
 
423
  # Error geometry
424
 
425
- HPC can also score candidate errors using the D₆ Vesica decomposition.
426
 
427
- For paired error components:
428
-
429
- $$
430
- v_p=e_p+e_{p+h}
431
- $$
432
 
433
- and
434
 
435
  $$
436
- w_p=e_p-e_{p+h},
437
  $$
438
 
439
- the error is separated into:
440
-
441
- * **Vesica / DC-like error**
442
- * **wave / AC-like error**
443
-
444
- The weighted objective is:
445
-
446
  $$
447
- E=
448
- \frac12
449
- \left(
450
- 4E_{\mathrm{vesica}}
451
- +
452
- E_{\mathrm{wave}}
453
- \right).
454
  $$
455
 
456
- Thus:
457
 
458
- $$
459
- \boxed{
460
- E_{\mathrm{vesica}}
461
- \text{ receives 4Γ— the penalty of }
462
- E_{\mathrm{wave}}.
463
- }
464
- $$
465
 
466
- The goal is not simply to minimize Euclidean distance. HPC attempts to distinguish error modes that behave differently during downstream computation.
467
 
468
  ---
469
 
470
- # Precision strategy
471
 
472
- HPC is designed for aggressive low-bit quantization and does not require every tensor to be treated identically.
473
 
474
- A typical strategy is:
475
 
476
- | Tensor class | HPC strategy |
477
- | ------------------------- | ----------------------- |
478
- | Attention projections | Q4_0 / higher precision |
479
- | FFN / MLP weights | Q2_K |
480
- | Expert weights | Q2_K |
481
- | Embeddings | Preserved / promoted |
482
- | Norms | Preserved |
483
- | Router / gate parameters | Preserved or promoted |
484
- | LM head / tied embeddings | Promoted when required |
485
 
486
- The exact routing is model-dependent.
487
 
488
- The purpose is straightforward:
 
 
489
 
490
- > Spend precision where small perturbations are most expensive, and aggressively compress where the model has greater tolerance.
491
 
492
  ---
493
 
494
- # Calibration
495
 
496
- HPC supports importance-matrix calibration for aggressive quantization.
497
 
498
- A typical workflow is:
499
 
500
- ```bash
501
- python3 LLM/generate_imatrix.py \
502
- model.gguf \
503
- calibration_data.txt \
504
- -o imatrix.dat \
505
- --chunks 5 \
506
- --verbose
 
 
507
  ```
508
 
509
- The resulting importance data can then be supplied to the re-quantizer.
510
 
511
  ---
512
 
513
- # Building
514
 
515
- Install the required dependencies on Ubuntu/Debian:
516
 
517
- ```bash
518
- sudo apt install \
519
- gcc \
520
- libgmp-dev \
521
- libmpfr-dev \
522
- python3 \
523
- python3-numpy
524
- ```
525
 
526
- Build the HPC engine:
527
 
528
- ```bash
529
- make -f makefile.quantize
530
- ```
531
 
532
- The build should produce:
 
 
 
 
533
 
534
- ```text
535
- libhexstate_q2k.so
536
- ```
537
 
538
  ---
539
 
540
- # Quantization
541
 
542
- Convert the source model to BF16 GGUF using llama.cpp:
543
 
544
- ```bash
545
- python3 llama.cpp/convert_hf_to_gguf.py \
546
- /path/to/model \
547
- --outfile Model-BF16.gguf \
548
- --outtype bf16
 
 
 
 
 
 
549
  ```
550
 
551
- Generate calibration data:
552
 
553
- ```bash
554
- python3 LLM/generate_imatrix.py \
555
- Model-BF16.gguf \
556
- calibration_data.txt \
557
- -o imatrix.dat \
558
- --chunks 5 \
559
- --verbose
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
560
  ```
561
 
562
- Run HPC quantization:
563
 
564
- ```bash
565
- python3 hexstate_requantize.py \
566
- Model-BF16.gguf \
567
- Model-Q2_K-HPC.gguf \
568
- --keep-metadata \
569
- --imatrix imatrix.gguf
570
- ```
571
 
572
  ---
573
 
574
- # Historical architecture
575
 
576
- HPC originally used a quantum-inspired implementation based on:
577
 
578
- * Z₆ state encoding
579
- * complex amplitudes
580
- * graph coupling
581
- * phase operations
582
- * IDFT₆
583
- * Griffiths–Niu-style sequential measurement
584
- * collapse/back-action
585
- * beam search
586
 
587
- That architecture was useful for exploring the underlying hypothesis:
588
 
589
- > **quantization states can be treated as interacting discrete variables rather than independent rounding decisions.**
590
 
591
- The current Sieve architecture keeps the useful structural concepts while replacing the Fourier/phase machinery with explicit probability scoring and dynamic programming.
592
 
593
- In other words:
594
 
595
- ```text
596
- Original HPC
597
-
598
- error
599
- ↓
600
- complex amplitudes
601
- ↓
602
- phase graph
603
- ↓
604
- IDFT₆
605
- ↓
606
- measurement
607
- ↓
608
- back-action
609
 
 
 
 
 
 
 
 
 
 
610
 
611
- Current HPC
612
 
613
- error
614
- ↓
615
- probability
616
- ↓
617
- Sieve
618
- ↓
619
- bounded back-action
620
- ↓
621
- 36-state lattice
622
- ↓
623
- Viterbi
624
- ```
625
 
626
- The latter is substantially easier to analyze, reproduce, and tune.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
627
 
628
  ---
629
 
630
- # Design philosophy
631
 
632
- HPC is built around several principles.
633
 
634
- ### 1. Do not assume minimum MSE is the whole objective
635
 
636
- A quantization configuration with slightly higher local error can still be preferable if its error structure is better behaved globally.
 
 
 
 
 
 
 
 
 
 
 
637
 
638
- ### 2. Preserve alternatives
639
 
640
- Low-bit quantization has a highly discrete error landscape. Prematurely selecting one candidate can eliminate useful configurations.
 
 
 
 
 
 
 
 
 
 
 
 
 
641
 
642
- ### 3. Model interactions explicitly
643
 
644
- Quantization blocks do not exist in complete isolation during inference.
645
 
646
- HPC therefore gives neighboring state assignments an explicit influence.
647
 
648
- ### 4. Encourage diversity without forcing it
649
 
650
- The Sieve's 0.85 penalty is deliberately bounded.
651
 
652
- It discourages repeated state assignments but does not prohibit them.
 
 
653
 
654
- ### 5. Separate local inference from global optimization
655
 
656
- The Sieve supplies local state information.
657
 
658
- Viterbi then searches the resulting 36-state Q2 lattice globally.
 
 
 
 
 
 
 
 
659
 
660
- ### 6. Keep a reconstruction-quality escape hatch
661
 
662
- The 5% greedy override prevents the global model from making obviously bad local decisions.
 
 
 
 
663
 
664
  ---
665
 
666
- # What HPC is β€” and isn't
667
 
668
- HPC is:
669
 
670
- * an experimental ultra-low-bit quantization engine
671
- * MIT licensed
672
- * designed for GGUF/llama.cpp workflows
673
- * focused on Q2-class compression
674
- * based on structured discrete-state optimization
675
- * capable of combining local reconstruction error with state interactions
676
 
677
- HPC is **not**:
678
 
679
- * a claim that LLM inference is literally quantum computation
680
- * a replacement for llama.cpp itself
681
- * a guarantee of improved perplexity on every model
682
- * a conventional round-to-nearest quantizer
683
- * dependent on quantum hardware
684
 
685
- The current Sieve is entirely classical.
686
 
687
- The "quantum-inspired" terminology refers to the historical mathematical lineage and the state/graph representationβ€”not to a requirement for quantum hardware.
688
 
689
  ---
690
 
@@ -692,56 +699,38 @@ The "quantum-inspired" terminology refers to the historical mathematical lineage
692
 
693
  HPC-Quantize is released under the **MIT License**.
694
 
695
- Quantized model files remain subject to the license and terms of the underlying base model.
 
 
696
 
697
  ---
698
 
699
  # Status
700
 
701
- HPC-Quantize is research software.
702
 
703
- The quantization architecture is intentionally experimental, particularly at extremely low bitrates such as Q2.
704
 
705
- Results should therefore be evaluated using:
706
 
707
- * perplexity
708
- * task benchmarks
709
- * generation quality
710
- * long-context behavior
711
- * mathematical reasoning
712
- * coding benchmarks
713
- * model-specific evaluations
714
 
715
- rather than RMSE alone.
716
 
717
  ---
718
 
719
- ## The short version
720
 
721
- Conventional Q2:
722
 
723
- $$
724
- \boxed{
725
- \text{pick the lowest-error candidate}
726
- }
727
- $$
728
 
729
- HPC:
730
 
731
- $$
732
- \boxed{
733
- \text{score candidates}
734
- \rightarrow
735
- \text{encode state}
736
- \rightarrow
737
- \text{sieve}
738
- \rightarrow
739
- \text{condition neighbors}
740
- \rightarrow
741
- \text{search the 36-state lattice}
742
- \rightarrow
743
- \text{select a globally coherent configuration}
744
- }
745
- $$
746
 
747
- **HPC-Quantize treats ultra-low-bit quantization as a structured state-selection problem, not merely a rounding problem.**
 
2
  license: mit
3
  ---
4
 
 
 
 
 
 
 
 
 
 
 
 
5
  # HPC-Quantize
6
 
7
+ ## Holographic Phase Contraction for Ultra-Low-Bit LLM Quantization
8
 
9
+ **HPC-Quantize** is an experimental, MIT-licensed quantization engine for aggressively compressing large language models into extremely low-bit formats, with a particular focus on **Q2-class quantization**.
10
 
11
+ The central idea is simple:
12
 
13
+ > **At very low bitrates, quantization should be treated as a structured reconstruction problem rather than independent rounding of individual blocks.**
14
 
15
+ Instead of choosing a quantization candidate solely from its local reconstruction error, HPC generates competing reconstructions, represents them in a compact discrete state space, models interactions between neighboring blocks, and performs a global sequence optimization before writing the final GGUF.
16
 
17
+ The current production path is entirely classical. Earlier versions explored quantum-inspired state and measurement formulations; the current implementation uses a **sequential Sieve, bounded state back-action, and a 36-state Viterbi optimizer** for Q2.
18
 
19
  ---
20
 
21
+ ## Why HPC?
22
 
23
+ At Q4 or Q5, a model often has enough representational freedom that many quantization strategies work reasonably well.
24
 
25
+ At Q2, the situation changes dramatically.
 
 
 
 
 
 
26
 
27
+ A block has very few representable values, so small decisions about scale, minimum, and code assignment can produce disproportionately large changes in the resulting weight tensor.
28
 
29
+ A conventional quantizer often reduces the problem to:
 
 
 
 
 
 
30
 
31
+ ```text
32
+ original weights
33
+ β”‚
34
+ β–Ό
35
+ find locally best parameters
36
+ β”‚
37
+ β–Ό
38
+ encode quantized block
39
+ ```
40
 
41
+ HPC instead treats each block as a **discrete candidate-selection problem**:
42
 
43
  ```text
44
+ original weights
45
+ β”‚
46
+ β–Ό
47
+ generate candidate reconstructions
48
+ β”‚
49
+ β–Ό
50
+ score candidate errors
51
+ β”‚
52
+ β–Ό
53
+ map candidates into a compact state space
54
+ β”‚
55
+ β–Ό
56
+ sequential Sieve
57
+ β”‚
58
+ β–Ό
59
+ Q2 state lattice
60
+ β”‚
61
+ β–Ό
62
+ global Viterbi optimization
63
+ β”‚
64
+ β–Ό
65
+ select physical reconstructions
66
+ β”‚
67
+ β–Ό
68
+ GGUF
 
 
 
 
 
69
  ```
70
 
71
+ This lets the optimizer preserve competing possibilities until there is enough information to make a global decision.
72
+
73
  ---
74
 
75
+ # What HPC actually does
76
 
77
+ HPC is still a **quantizer/re-quantizer**.
78
 
79
+ It does not retrain the neural network, modify the model architecture, or learn a new set of representations.
 
 
80
 
81
+ The "reconstruction" terminology refers to what happens during candidate selection: HPC explicitly constructs multiple possible low-bit approximations of the original weights and evaluates them against the source weights.
82
 
83
+ Conceptually:
84
 
85
+ $$
86
+ W \rightarrow
87
+ \{\hat W_1,\hat W_2,\ldots,\hat W_n\}
88
+ \rightarrow
89
+ \text{structured candidate selection}
90
+ \rightarrow
91
+ Q(W)
92
+ $$
93
+
94
+ The final output remains a normal quantized GGUF model.
95
 
96
  ---
97
 
98
+ # Core design
99
 
100
+ HPC combines several ideas:
101
 
102
+ * candidate reconstruction
103
+ * weighted reconstruction error
104
+ * optional importance-matrix weighting
105
+ * discrete state mapping
106
+ * sequential Sieve selection
107
+ * bounded neighboring-state back-action
108
+ * Q2 state coupling
109
+ * global Viterbi optimization
110
+ * local reconstruction-quality safeguards
111
 
112
+ The important distinction is that these components operate **together** rather than treating every quantization block as completely independent.
113
 
114
+ ---
115
 
116
+ # Candidate generation
117
+
118
+ For an eligible Q2 block, HPC searches over possible quantization parameters rather than committing immediately to a single local solution.
119
 
120
+ For Q2_K, the important coupled parameters are represented conceptually as:
121
 
122
  $$
123
+ (d,d_{\min})
124
  $$
125
 
126
+ Candidate parameters are evaluated by reconstructing the quantized block and measuring its error against the original weights.
127
+
128
+ With an importance matrix, the error can be weighted so that sensitive dimensions contribute more heavily:
129
 
130
  $$
131
+ E =
132
+ \sum_i w_i(x_i-\hat{x}_i)^2.
133
  $$
134
 
135
+ This is useful because ordinary unweighted RMSE assumes every weight contributes equally to the final model behavior.
136
 
137
+ HPC can therefore evaluate:
 
 
138
 
139
+ > **How well does this candidate reconstruct the important parts of the original block?**
140
 
141
+ rather than only:
 
 
 
 
 
 
 
 
 
 
142
 
143
+ > **How small is its raw Euclidean error?**
144
 
145
  ---
146
 
147
+ # From candidates to states
148
 
149
+ Keeping every physical candidate in the global optimizer would be expensive.
150
 
151
+ HPC therefore maps candidates into a compact symbolic state representation.
 
 
 
 
 
152
 
153
+ The current implementation uses **six symbolic states** for each quantization parameter:
154
 
155
  $$
156
+ d \in \{0,1,2,3,4,5\}.
 
 
 
157
  $$
158
 
159
+ The symbolic states provide a compact representation of the candidate landscape.
160
 
161
+ Multiple physical candidates can belong to the same symbolic state.
162
 
163
+ This separation is important:
164
 
165
+ ```text
166
+ physical candidates
167
+ β”‚
168
+ β–Ό
169
+ symbolic state representation
170
+ β”‚
171
+ β–Ό
172
+ global optimization
173
+ β”‚
174
+ β–Ό
175
+ physical candidate selection
176
+ ```
177
 
178
+ The symbolic state space is therefore not the same thing as the number of physical reconstructions generated during candidate search.
179
 
180
+ ---
 
 
181
 
182
+ # The Q2 state space
183
 
184
+ Q2 has two coupled parameters:
185
 
186
  $$
187
+ (d,d_{\min}).
 
 
 
 
 
 
188
  $$
189
 
190
+ Each is represented by six symbolic states.
191
+
192
+ Therefore the joint Q2 space contains:
193
 
194
  $$
195
+ 6\times6=36
196
  $$
197
 
198
+ states.
199
+
200
+ The state index is:
201
 
202
  $$
203
+ s=6q_D+q_M
 
 
 
204
  $$
205
 
206
+ where:
207
 
208
  $$
209
+ q_D,q_M\in\{0,\ldots,5\}.
 
 
 
 
210
  $$
211
 
212
+ The resulting lattice is:
213
 
214
+ ```text
215
+ dmin
216
+ 0 1 2 3 4 5
217
+ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€
218
+ 0 β”‚ 0 1 2 3 4 5
219
+ 1 β”‚ 6 7 8 9 10 11
220
+ d 2 β”‚ 12 13 14 15 16 17
221
+ 3 β”‚ 18 19 20 21 22 23
222
+ 4 β”‚ 24 25 26 27 28 29
223
+ 5 β”‚ 30 31 32 33 34 35
224
+ ```
225
 
226
+ The 36-state representation is a **global optimization space**, not a claim that only 36 physical quantization candidates exist.
 
 
227
 
228
+ ---
229
 
230
+ # Candidate probabilities
231
 
232
+ Rather than treating each candidate as simply "best" or "not best", HPC can preserve information about the relative quality of competing candidates.
 
 
233
 
234
+ Candidate errors are converted into probability-like weights using a Boltzmann-style transformation:
235
 
236
  $$
237
+ P_i \propto e^{-T(E_i-E_{\min})}.
238
  $$
239
 
240
+ This creates a soft candidate distribution.
241
 
242
+ The practical consequence is important:
 
 
243
 
244
+ > A candidate that is slightly worse than the local minimum can remain relevant if it belongs to a useful region of the state space.
245
 
246
+ That makes the search less eager to collapse immediately to a single local minimum.
 
 
247
 
248
  ---
249
 
250
+ # The Sieve
251
 
252
+ The current production architecture uses a **sequential Sieve**.
253
 
254
+ Rather than treating blocks as fully independent, neighboring state distributions influence one another.
 
 
255
 
256
+ For a candidate state \(d\), the implementation applies a bounded compatibility penalty based on neighboring probability:
257
 
258
  $$
259
+ B_j(d)=1-0.15p_j(d).
260
  $$
261
 
262
+ The combined Sieve score can be written conceptually as:
263
 
264
  $$
265
+ S(d)=p_0(d)\prod_j(1-0.15p_j(d)).
 
 
266
  $$
267
 
268
+ The effect is deliberately modest.
269
 
270
+ If a neighboring site is strongly concentrated on the same state, that state becomes somewhat less attractive locally.
271
 
272
+ The Sieve therefore encourages **state diversity and compatibility** without forcing an alternating pattern.
273
 
274
  ---
275
 
276
+ # Sieve slack
277
 
278
+ The Sieve also avoids immediately discarding every candidate that is not the top local state.
279
 
280
+ A bounded slack region allows near-optimal states to survive the first selection stage.
 
 
281
 
282
+ Conceptually:
283
 
284
+ ```text
285
+ best state
286
+ β”‚
287
+ β”œβ”€β”€ keep
288
+ β”‚
289
+ β”œβ”€β”€ keep near-optimal alternatives
290
+ β”‚
291
+ └── discard clearly inferior states
292
+ ```
293
 
294
+ This prevents the candidate distribution from collapsing too early.
295
 
296
+ ---
297
+
298
+ # Sequential conditioning
299
+
300
+ Once the current state is selected, its influence is propagated into neighboring sites.
301
 
302
+ The selected state therefore affects subsequent decisions.
303
 
304
+ This gives the process a sequential character:
305
 
306
  $$
307
  P(d_1)
308
  \rightarrow
309
+ P(d_2|d_1)
310
  \rightarrow
311
+ P(d_3|d_1,d_2)
312
+ \rightarrow \cdots
313
  $$
314
 
315
+ The quantization process is consequently no longer just a collection of independent block decisions.
 
 
316
 
317
  ---
318
 
319
+ # Viterbi optimization
320
 
321
+ After the Sieve, Q2 state probabilities are expanded into the full **36-state joint lattice**.
322
 
323
+ Each block receives a local cost combining reconstruction error with state probability.
324
 
325
+ Conceptually:
 
 
 
 
 
 
326
 
327
  $$
328
+ C_i(s)
329
+ =
330
+ E_i(s)-\lambda \log P_i(s).
331
  $$
332
 
333
+ HPC then adds a transition cost between neighboring blocks.
334
 
335
+ A simple form is:
 
 
 
 
 
 
 
336
 
337
  $$
338
+ T(s',s)
339
+ \propto
340
+ |q_D-q'_D|+
341
+ |q_M-q'_M|.
 
 
 
342
  $$
343
 
344
+ This is Manhattan distance on the 6Γ—6 state lattice.
345
 
346
+ The dynamic-programming recurrence is:
347
 
348
  $$
349
+ DP_i(s)
350
+ =
351
+ C_i(s)
352
+ +
353
  \min_{s'}
354
  \left[
355
  DP_{i-1}(s')+T(s',s)
356
  \right].
357
  $$
358
 
359
+ The result is a **globally optimized sequence of Q2 states**.
360
+
361
+ This is one of the major differences between HPC and purely local quantization.
362
 
363
  ---
364
 
365
+ # Local safety
366
 
367
+ Global regularization should not be allowed to produce obviously poor local reconstructions.
368
 
369
+ After global selection, HPC can compare the chosen state against the locally best reconstruction.
370
 
371
+ A sufficiently large local improvement can trigger a local override.
372
 
373
+ This gives the optimizer a safety mechanism:
 
 
 
 
374
 
375
+ ```text
376
+ global structure
377
+ β”‚
378
+ β–Ό
379
+ candidate selected
380
+ β”‚
381
+ β–Ό
382
+ is the local reconstruction much better?
383
+ β”‚
384
+ β”Œβ”€β”€β”΄β”€β”€β”
385
+ β”‚ β”‚
386
+ yes no
387
+ β”‚ β”‚
388
+ local keep
389
+ winner global
390
+ winner
391
+ ```
392
 
393
+ The goal is to prevent the global objective from becoming disconnected from actual reconstruction quality.
394
 
395
  ---
396
 
397
  # Error geometry
398
 
399
+ HPC can also use structured error decomposition rather than treating every error component identically.
400
 
401
+ One experimental component uses a D₆/Vesica-style decomposition of paired error terms.
 
 
 
 
402
 
403
+ For a pair of error components:
404
 
405
  $$
406
+ v=e_p+e_{p+h}
407
  $$
408
 
 
 
 
 
 
 
 
409
  $$
410
+ w=e_p-e_{p+h}.
 
 
 
 
 
 
411
  $$
412
 
413
+ This separates the error into different modes before applying the final weighting.
414
 
415
+ The intent is to distinguish error geometry rather than assuming that all directions in weight space have identical consequences.
 
 
 
 
 
 
416
 
417
+ This is an experimental feature of the HPC objective rather than a requirement of GGUF or Q2_K itself.
418
 
419
  ---
420
 
421
+ # Building
422
 
423
+ HPC-Quantize is intended to be used alongside a GGUF/llama.cpp workflow.
424
 
425
+ Typical dependencies include:
426
 
427
+ ```bash
428
+ sudo apt install \
429
+ gcc \
430
+ libgmp-dev \
431
+ libmpfr-dev \
432
+ python3 \
433
+ python3-numpy
434
+ ```
 
435
 
436
+ Build the native quantization component:
437
 
438
+ ```bash
439
+ make -f makefile.quantize
440
+ ```
441
 
442
+ The resulting library/binary names may vary with the current revision.
443
 
444
  ---
445
 
446
+ # Mixed-precision workflows
447
 
448
+ HPC is primarily intended to solve the problem of **aggressive compression**, not to force every tensor in a model into identical precision.
449
 
450
+ A practical deployment may therefore retain higher precision for particularly sensitive tensors and use Q2 for the bulk of the model.
451
 
452
+ For example:
453
+
454
+ ```text
455
+ Model
456
+ β”œβ”€β”€ embeddings β†’ higher precision
457
+ β”œβ”€β”€ normalization β†’ preserved
458
+ β”œβ”€β”€ attention β†’ Q4 / promoted
459
+ β”œβ”€β”€ FFN / experts β†’ Q2
460
+ └── other large mats β†’ Q2
461
  ```
462
 
463
+ The optimal allocation is model-dependent.
464
 
465
  ---
466
 
467
+ # Why Q2?
468
 
469
+ Q2 is where conventional quantization becomes particularly unforgiving.
470
 
471
+ At higher precision, the quantizer has many representational degrees of freedom.
 
 
 
 
 
 
 
472
 
473
+ At Q2, many distinct original weight values must share a very small set of reconstruction values.
474
 
475
+ This means:
 
 
476
 
477
+ $$
478
+ \text{small parameter change}
479
+ \rightarrow
480
+ \text{large discrete reconstruction change}.
481
+ $$
482
 
483
+ HPC is designed around this regime.
484
+
485
+ Rather than assuming the locally nearest reconstruction is always globally best, it explicitly searches among competing discrete configurations.
486
 
487
  ---
488
 
489
+ # HPC versus conventional quantization
490
 
491
+ A simplified conventional pipeline is:
492
 
493
+ ```text
494
+ weight block
495
+ β”‚
496
+ β–Ό
497
+ estimate scale/minimum
498
+ β”‚
499
+ β–Ό
500
+ round values
501
+ β”‚
502
+ β–Ό
503
+ write block
504
  ```
505
 
506
+ A simplified HPC pipeline is:
507
 
508
+ ```text
509
+ weight block
510
+ β”‚
511
+ β–Ό
512
+ generate competing reconstructions
513
+ β”‚
514
+ β–Ό
515
+ score candidates
516
+ β”‚
517
+ β–Ό
518
+ map to symbolic states
519
+ β”‚
520
+ β–Ό
521
+ Sieve + neighboring interaction
522
+ β”‚
523
+ β–Ό
524
+ construct Q2 state lattice
525
+ β”‚
526
+ β–Ό
527
+ Viterbi global optimization
528
+ β”‚
529
+ β–Ό
530
+ select physical reconstructions
531
+ β”‚
532
+ β–Ό
533
+ write Q2_K
534
  ```
535
 
536
+ The difference is not that HPC stops being quantization.
537
 
538
+ The difference is **how much structure it retains before committing to the final quantized representation**.
 
 
 
 
 
 
539
 
540
  ---
541
 
542
+ # A useful way to think about HPC
543
 
544
+ HPC can be viewed as three nested optimization problems:
545
 
546
+ ### Local reconstruction
 
 
 
 
 
 
 
547
 
548
+ > Which low-bit approximation best represents this block?
549
 
550
+ ### State inference
551
 
552
+ > Which region of the discrete candidate space is promising?
553
 
554
+ ### Global sequence optimization
555
 
556
+ > Which sequence of candidate states produces the best overall configuration?
557
+
558
+ That can be summarized as:
 
 
 
 
 
 
 
 
 
 
 
559
 
560
+ $$
561
+ \boxed{
562
+ \text{reconstruction}
563
+ +
564
+ \text{state inference}
565
+ +
566
+ \text{global optimization}
567
+ }
568
+ $$
569
 
570
+ rather than:
571
 
572
+ $$
573
+ \boxed{
574
+ \text{independent rounding}
575
+ }
576
+ $$
 
 
 
 
 
 
 
577
 
578
+ ---
579
+
580
+ # Experimental nature
581
+
582
+ HPC is research software.
583
+
584
+ It should not be assumed that:
585
+
586
+ * lower RMSE always produces better model behavior;
587
+ * lower perplexity always produces better reasoning;
588
+ * one quantization strategy wins on every architecture;
589
+ * Q2 quality transfers perfectly between models;
590
+ * state-interaction parameters are universally optimal.
591
+
592
+ The correct way to evaluate HPC is with a combination of:
593
+
594
+ * reconstruction error
595
+ * perplexity
596
+ * reasoning benchmarks
597
+ * mathematical evaluation
598
+ * coding tasks
599
+ * long-context tests
600
+ * instruction following
601
+ * qualitative generation
602
+ * memory usage
603
+ * inference speed
604
+
605
+ The objective of HPC is not to optimize one number in isolation.
606
 
607
  ---
608
 
609
+ # Current architecture
610
 
611
+ The project originally explored a more explicitly quantum-inspired formulation involving state amplitudes, phase operations, graph coupling, Fourier/IDFT transforms, and sequential measurement.
612
 
613
+ The current engine has moved toward a more explicit classical formulation:
614
 
615
+ ```text
616
+ Historical approach
617
+ ───────────────────
618
+ candidate error
619
+ ↓
620
+ amplitudes / phase
621
+ ↓
622
+ graph coupling
623
+ ↓
624
+ measurement
625
+ ↓
626
+ back-action
627
 
 
628
 
629
+ Current approach
630
+ ────────────────
631
+ candidate error
632
+ ↓
633
+ probability distribution
634
+ ↓
635
+ sequential Sieve
636
+ ↓
637
+ bounded back-action
638
+ ↓
639
+ 36-state Q2 lattice
640
+ ↓
641
+ Viterbi
642
+ ```
643
 
644
+ The mathematical intuition of interacting discrete states remains, but the current production implementation is classical and deterministic.
645
 
646
+ ---
647
 
648
+ # Future directions
649
 
650
+ The state lattice is intentionally compact.
651
 
652
+ The 36-state Q2 space is:
653
 
654
+ $$
655
+ 6\times6=36.
656
+ $$
657
 
658
+ That does **not** mean the physical candidate space must contain only 36 candidates.
659
 
660
+ Possible future work includes:
661
 
662
+ * more symbolic states per parameter;
663
+ * multiple physical candidates retained per symbolic state;
664
+ * beam search inside individual states;
665
+ * hierarchical state refinement;
666
+ * adaptive state resolution;
667
+ * larger candidate beams for difficult tensors;
668
+ * tensor-dependent state cardinality;
669
+ * improved transition models;
670
+ * architecture-specific state priors.
671
 
672
+ One particularly interesting extension is to retain multiple physical candidates for each symbolic state:
673
 
674
+ $$
675
+ 36\times K.
676
+ $$
677
+
678
+ This would preserve more of the physical reconstruction landscape while keeping the coarse 6Γ—6 state structure.
679
 
680
  ---
681
 
682
+ # What HPC is trying to preserve
683
 
684
+ At ultra-low precision, numerical error is inevitable.
685
 
686
+ The goal is therefore not:
 
 
 
 
 
687
 
688
+ > **make every weight numerically perfect.**
689
 
690
+ The goal is:
 
 
 
 
691
 
692
+ > **spend the available representational capacity where it matters most, preserve competitive reconstruction alternatives long enough for global selection, and avoid treating every block as an isolated rounding problem.**
693
 
694
+ That is the central design philosophy of HPC-Quantize.
695
 
696
  ---
697
 
 
699
 
700
  HPC-Quantize is released under the **MIT License**.
701
 
702
+ The licensing terms of any model quantized with HPC remain separate from the HPC software license.
703
+
704
+ Always verify the license of the underlying base model before redistribution.
705
 
706
  ---
707
 
708
  # Status
709
 
710
+ **Experimental / research software**
711
 
712
+ The current engine is actively evolving, particularly around ultra-low-bit Q2 quantization.
713
 
714
+ The project currently prioritizes:
715
 
716
+ * low-bit reconstruction quality
717
+ * model coherence
718
+ * reasoning preservation
719
+ * structured state selection
720
+ * aggressive memory reduction
 
 
721
 
722
+ over compatibility with any single traditional quantization metric.
723
 
724
  ---
725
 
726
+ ## In one sentence
727
 
728
+ **HPC-Quantize is a structured ultra-low-bit quantizer that searches over competing Q2 reconstructions, reasons about them as interacting discrete states, and uses global sequence optimization to choose the final GGUF configuration.**
729
 
730
+ ---
 
 
 
 
731
 
732
+ ## Acknowledgements
733
 
734
+ HPC-Quantize builds on the broader GGUF and llama.cpp ecosystem and is intended to interoperate with existing llama.cpp-based tooling.
 
 
 
 
 
 
 
 
 
 
 
 
 
 
735
 
736
+ The project also explores ideas inspired by discrete graphical models, sequential inference, information-weighted reconstruction, and quantum-inspired state representations.