why only 13b active on the flash?
Qwen 27b eats this alive. 13b A is not enough. We need 50b at least
This is a flash version,why don't you consider the Pro Version?
Also ,Qwen3.5 only activates 17b params.
This is a flash version,why don't you consider the Pro Version?
Also ,Qwen3.5 only activates 17b params.
3.5 is old and the pro is huge wtf
yes exactly. who can even fit 1.6t? This is not optimized for local at all. Give us a 200b total 100b dense
Qwen 27b eats this alive. 13b A is not enough. We need 50b at least
yes, 13B is a way too small. we need at least 37B
Qwen 27b eats this alive. 13b A is not enough. We need 50b at least
Within a ~248k context window Qwen 27B is good but with context scaled past that is not very stable in my experience. For 1M context window, DeepSeek V4 Flash better option for hardware constrained 'prosumers'