Please collaborate with PrismML on a Bonsai version

#17
by mindplay - opened

At 33B, a 1-bit version would likely be usable on a 16 GB GPU, and a 1.5-bit version for 24 GB GPUs.

Their approach is not as simple as just quantizing - post training is required to recover precision, so you would probably need to collaborate with them.

The community will likely quantize this model to 4-bit, but it's going to be so much worse than what PrismML can do with 1-bit.

A Bonsai version of this model would be light-years ahead of anything else that runs locally today.

Sign up or log in to comment