Wow this is an outright killer model... anyone can now run this with mere 48 gb ram/vram & a fast nvme drive

#30
by mayankiit04 - opened

Wow this is an outright killer model... anyone can now run this with mere 32 gb ram & a fast nvme drive ... darn cheap

I gonna ask question I have 32gb ram 16gb vram and fast nvme ssd can i run this model decent speed.

yes, you will need to push -ngl 48 and try with -ncmoe with 40s and lower quant model of 1/2 bit... if total file size is great then 48 gb as it is in my case you need for now --mmap too. TPS is generally good in 20s... prompt is bit slow in 40s-80s for me ... few folks are already working to ensure that like ngl if we can push experts onto ram, leaving ngram for mmap ... that will boost up prompt prefill over 400 ...

Wow this is an outright killer model... anyone can now run this with mere 32 gb ram & a fast nvme drive ... darn cheap

I’d like to ask: following the author's description, how do I run this using vLLM via Docker? Could you provide the specific parameters?

mayankiit04 changed discussion title from Wow this is an outright killer model... anyone can now run this with mere 32 gb ram & a fast nvme drive to Wow this is an outright killer model... anyone can now run this with mere 48 gb ram/vram & a fast nvme drive

Sign up or log in to comment