somersetent 's Collections

Quantization/Quantized

Quantization in LLMs compresses model weights from high-precision formats (like 16-bit) to lower-precision formats (like 8-bit or 4-bit).