somersetent 's Collections

GPTQ

GPTQ is a fast, accurate post-training quantization method for LLMs. It compresses 16-bit model weights down to 4-bit integers.