somersetent 's Collections

5-bit Quantization

5-bit quantization in LLMs means compressing the model's weights from high precision (like 16-bit) down to an average of 5 bits per weight.