35B-A3B or 35B-A5B

#3
by MaxDaddyLongs - opened

35B-A3B or 35B-A5B
Humanity needs this MoE model...

yes cant run 180B param model

We really need that model. The % of people that own machines with unified memory or high end gpus for AI like 3090, rtx 4080 and higher is not that high at all, they are really low % of the pc community.

Yeah, we’re all really waiting for this model because it’s the best for personal pc

Cant we just offload the ngram params (and maybe also other inactive params) to SSD? I am quite optimistic that this can actually run and run well on smaller machines

I agree, it is what we need Qwen for. A3B, maybe A5B, is absolutely mandatory.

PLS ADD 36B-4B

MoE is the best!

True, We the people need a 35-40b moe model🥺

Qwen 3.6 35b a3b is an incredible model for this kind of hardware. tuned and quantized i can hit 30-45 tokens a second. It's a real need.

whatever you want, only keep in mind that fit in 24gb vram, please

Crossing my fingers for this. I'm running an AMD HX370 with 32GB ram, and the 3.6 35B moe has been perfect for it. I get 28-32 tps and has 128K context. So just a dev step or two more and this will cover 95% of all my AI needs.

Sign up or log in to comment