Average decode token number only 1.6

#4
by wc-llm - opened

I test this model with sglang, the average decode token number only is 1.6.It only increased the speed by 10%

Seems you only see real speedup with enable thinking on.

Sign up or log in to comment