somersetent 's Collections

2-bit Quantization

2-bit quantization reduces the memory required to store each model weight from 16 bits down to 2 bits. This compresses the model size by roughly 8x.