GLM 5.3 Flash Lite idea

#13
by ChessVania - opened

24b, 12b active (because 3b is fast, but bad)

alternatively, 60B-a6b

12b active is too much for a 24b parameters model.
5-6b would be the ideal.

alternatively, 60B-a6b

too big for most GPU's

12b active is too much for a 24b parameters model.
5-6b would be the ideal.

But much smarter than 6b

alternatively, 60B-a6b

too big for most GPU's

Best use of MoE models is to offload the experts to RAM. The active 6B parameters can fit in a 8GB GPU as long as you have enough RAM for the rest.

Sign up or log in to comment