File size: 1,951 Bytes
762c004
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
---
title: Quantization Explorer
emoji: ⚙️
colorFrom: blue
colorTo: indigo
sdk: static
pinned: false
license: mit
short_description: "Explore quantization: FP8, INT8, INT4 and trade-offs."
---

# Quantization Explorer

**Quantization Explorer** is an educational Hugging Face Space by the [`open-weight`](https://huggingface.co/open-weight) organization.

It explains how model quantization reduces memory requirements by representing weights at lower precision, and how common approaches such as **FP8, INT8, INT4, bitsandbytes, GPTQ, AWQ and GGUF quantization** differ in purpose and trade-offs.

## What you can explore

- What model quantization is
- FP16/BF16 vs. FP8 vs. INT8 vs. INT4
- Theoretical raw weight memory
- Post-training quantization
- On-the-fly quantization
- Calibration-based methods
- bitsandbytes
- GPTQ
- AWQ
- GGUF / llama.cpp quantization
- Quality, speed and compatibility trade-offs
- A simple quantization decision helper

## Core idea

```text
Higher-precision weights
        ↓
Quantization method
        ↓
Lower-bit representation
        ↓
Lower memory / storage
        ↓
Potential speed benefits
        +
Possible quality / compatibility trade-offs
```

## Primary references

- Hugging Face Transformers — Quantization overview: https://huggingface.co/docs/transformers/quantization/overview
- bitsandbytes: https://huggingface.co/docs/transformers/en/quantization/bitsandbytes
- GPTQ: https://huggingface.co/docs/transformers/quantization/gptq
- AWQ: https://huggingface.co/docs/transformers/quantization/awq
- llama.cpp quantization: https://github.com/ggml-org/llama.cpp/tree/master/tools/quantize

## Related organization

Open Weight  
https://huggingface.co/open-weight

## Related project

Open Weights  
https://huggingface.co/open-weights

## Collaboration

Open-weight AI, model infrastructure, inference, deployment, research and ecosystem partnerships.

**Contact:** agenten@magenta.de