--- pipeline_tag: text-generation license: other license_name: modified-mit license_link: https://github.com/MiniMax-AI/MiniMax-M2.7/blob/main/LICENSE library_name: mlx base_model: MiniMaxAI/MiniMax-M2.7 tags: - mlx --- Even as of July 2026 on M3 Mac 512Gb RAM I can't seem to find something better than Minimax 2.7, so here's the OMLX 8-bit oQ8e optimized quant, on paper it should be better than a regular 8-bit quant available elsewhere on HF. The optimized 8-bit quant runs quite fast and with turboquant you can pretty easily fit other models on a 512Gb or at least get a huge KV cache size while gaining up to 20 TPS vs 15 on the unquantized model; the difference between annoying and quite practical. Below is the original base model card... # mlx-community/MiniMax-M2.7 This model [mlx-community/MiniMax-M2.7](https://huggingface.co/mlx-community/MiniMax-M2.7) was converted to MLX format from [MiniMaxAI/MiniMax-M2.7](https://huggingface.co/MiniMaxAI/MiniMax-M2.7) using mlx-lm version **0.31.3**. ## Use with mlx ```bash pip install mlx-lm ``` ```python from mlx_lm import load, generate model, tokenizer = load("mlx-community/MiniMax-M2.7") prompt = "hello" if tokenizer.chat_template is not None: messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_dict=False, ) response = generate(model, tokenizer, prompt=prompt, verbose=True) ```