Is it possible to make less than 1 bit quantization?

#1
by RealBar - opened

I'm look for if there is any possible methods to make large frontier models 10x, 20x smaller size, maybe some weights fusion techs?

just don't install it atp

@RealBar you can prune the model

I have tried some older GLM models froms Cerebras REAP (https://huggingface.co/collections/cerebras/cerebras-reap). They were pruned by about 20% and then being quantized (e.g. by Unsloth). But that is still not another 10-20x on top of quantization. REAPed models work ok, but at that point you're probably just chasing shadows.

There are plenty of good enough smaller models out there if you don't have a few spare millions of $ in the bank to whip-up terabytes of VRAM.

Unsloth AI org

That'll be hard - 1-bit is currently 86% smaller and retains around 76.2% accuracy

xD, man how much I want to see IQ0_XXXXXS but no, less then 1 bit quantization isn't possible with our today's compute. The tiniest unit in compute is a 1 or a 0 so... xD

unless if I have been lied to

You are crazy.

Yes, it is possible to do below-1-bit quantization, but it's not trivial to do. Basically, you have to pack individual values in tensors into groups and then quantize the groups - so you basically quantize something like a [0.5, 0.3, 1.2, 0.9] quadruple into say [-2]. As long as the bit-budget for the aggregate is smaller than the number of aggregates, you get a below-1-bit quant.

Just look at JPEG and MPEG. Less than one bit per element - pixel or weight - is possible, but not with quantization alone. You need a transformation on top. For images, DCT and friends work great. For NN weights there's no known transformation yet.
Yet.

This comment has been hidden (marked as Low Quality)

Yes, natively targeting a smaller model size is generally better.
But knowledge is not evenly distributed across model weights, so each model can be further compressed without affecting capabilities much. It's not lossless.

This is an interesting discussion. @TobDeBer 's reference to media lossy compression (JPEG and MPEG) is a very interesting twist. JPEG uses discrete cosine transforms that switches (x,y,color) coordinates to frequency values, the goal being reducing size by removing high frequency values efficiently. So that is a bit like what is already done with LLM quantization.

MPEG uses the video's time axis to describe changes with different types of frames. AFAIK, LLM inference works sequentially over each layer to produce a token, so a bit like playing back a video for each token. Maybe there is something to do there in regards to "compression".

You'll have to distill models at that point. You can imagine compressing below 1bit quant as deleting words from sentences. Your data would get all garbled up, leaving you with junk.

You may think of not using a transform. I fully agree in that case.
But just because nobody found a transform that works for weights doesn't mean it doesn't exist.
Think about what huge difference GIF with pure quantization is to JPEG with a good transform before quantization.
GIF looks bad at 3 bit depth no matter how much dithering you use.
JPEG still looks great at 0.2 bit.

Sign up or log in to comment