Sm120 support?

#1
by DopplerFreq - opened

Currently I cannot produce a vLLM recipe that allows this to be served on a 2x RTX PRO 6000 system. Specifically, I cannot get past the “Paged KV not supported on SM 12.0 in this PR” error. I am on the latest nightly build of vLLM in a docker Ubuntu environment. Is support for this quant on sm120+ planned for vLLM? I would love to try this model out.

Thinking Machines Lab org

I think that is more of a question for the vLLM folks. Might be worth asking under their repo

simon-thinky changed discussion status to closed

Sign up or log in to comment