8 bit quantized version using Intel's autoround. I had problems with 8 bit awq having weird outputs with tool calls, paths and variable names. This version work as expected.

This is the best option imo to run MiniMax-M2.7 at 8 bits (native precision) for Ampere cards, eg 3090's

Downloads last month
5
Safetensors
Model size
60B params
Tensor type
F32
I32
BF16
F16
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for bullerwins/MiniMax-M2.7-W8A16

Quantized
(118)
this model