The imatrix is GOAT!

#3
by sokann - opened

IQ2_KT with unsloth imatrix:

Final estimate: PPL over 565 chunks for n_ctx=512 = 3.8402 +/- 0.02168

IQ2_KT with muzzy imatrix:

Final estimate: PPL over 565 chunks for n_ctx=512 = 3.6907 +/- 0.02071

IQ1_S_R4 with unsloth imatrix:

Final estimate: PPL over 565 chunks for n_ctx=512 = 6.1435 +/- 0.03775

IQ1_S_R4 with muzzy imatrix:

Final estimate: PPL over 565 chunks for n_ctx=512 = 5.7371 +/- 0.03465

Thanks for the 5 days epic calculation ❤️

Wait.. when running llama-imatrix, were you using ubergarm-imatrix-calibration-corpus-v02.txt or wiki.test.raw?

Owner

Uh oh- wiki text raw is what I used. I didn’t realize there was the other one. I’ll start running against the ubergarm data set.

My misunderstanding, oops haha

Owner

I'm starting up the quant with ubergarm-imatrix-calibration-corpus-v02.txt. CPU only it's going to take me 6 days, I think I finally got my gpus involved though, so hopefully it will be faster. I will make a note on the model page.

Owner

This should be fixed now, thank you so much for noticing and correcting me!

Thanks for all your quants. Would you also be able to reupload the IQ1_KT quant please? It's in the same size group as IQ1_S_R4 but works better in my testing and I'd really like to try IQ1_KT with the new imatrix.

Owner

yes it's uploading now! my upload speed is super slow but it should be up in a few hours

Sign up or log in to comment