KV-cache quantization without any fork (recommended, 2026): upstream llama.cpp/Ollama now cover this natively β€” use -ctk q8_0 -ctv q8_0 (half KV memory, negligible quality loss: perplexity +0.002–0.05) or -ctk q4_0 -ctv q4_0 (quarter memory, β‰ˆ7.6% perplexity increase). In Ollama: OLLAMA_KV_CACHE_TYPE=q8_0 with OLLAMA_FLASH_ATTENTION=1. Keep K and V types symmetric to stay on the fast fused Flash-Attention path. Since April 2026, mainline llama.cpp also applies Hadamard rotation to KV activations (PR #21038), which greatly improves low-bit KV quality (opt-out: LLAMA_ATTN_ROT_DISABLE=1).

The RotorQuant/TurboQuant fork flow below is experimental/legacy: the TurboQuant llama.cpp PR was closed without merging (June 2026) and the fork is unmaintained relative to mainline. It is NOT required to use this model.

MERaLiON-2-10B-RotorQuant-MLX-4bit

MLX 4-bit RotorQuant quantization of aisingapore/MERaLiON-AudioLLM-Whisper-SEA-LION-V3-10B for Apple Silicon inference.

RotorQuant applies rotation-based quantization that decorrelates weight matrices before quantization, distributing outlier magnitudes more evenly across channels for improved accuracy at low bit-widths.

Model Specifications

Property Value
Base Model aisingapore/MERaLiON-AudioLLM-Whisper-SEA-LION-V3-10B
Parameters ~10B
Architecture Whisper encoder + Gemma-2-9B-IT decoder
Quantization RotorQuant 4-bit (MLX)
Disk Size ~5 GB
Peak RAM ~6 GB
License Apache 2.0
Task Automatic Speech Recognition / Speech-to-Text

Quickstart

Installation

pip install mlx-lm mlx-whisper

Inference

from mlx_lm import load, generate
from mlx_lm.cache import IsoQuantCache

model, tokenizer = load("majentik/MERaLiON-2-10B-RotorQuant-MLX-4bit")

# Create IsoQuantCache for RotorQuant models
cache = IsoQuantCache(model)

prompt = tokenizer.apply_chat_template(
    [{"role": "user", "content": "Transcribe the following audio."}],
    tokenize=False,
    add_generation_prompt=True,
)

response = generate(
    model,
    tokenizer,
    prompt=prompt,
    max_tokens=512,
    cache=cache,
)
print(response)

Quantization Details

RotorQuant is a rotation-based quantization strategy that:

  • Applies learned rotation matrices to decorrelate weight channels before quantization
  • Reduces the impact of outlier weights that typically degrade quantized model quality
  • Provides more uniform weight distributions, leading to better accuracy retention
  • Pairs with IsoQuantCache for consistent KV-cache quantization during inference

This 4-bit variant provides a strong balance between quality and memory usage. The rotation-based approach is particularly effective at 4-bit, where outlier sensitivity is more pronounced.

Supported Languages

MERaLiON-2 supports speech recognition in Southeast Asian languages including English, Mandarin Chinese, Malay, Tamil, and Indonesian.

Memory Estimates

Device Feasibility
MacBook Air M1 (8 GB) Feasible with limited headroom
MacBook Pro M1/M2 (16 GB) Comfortable
MacBook Pro M2/M3 (32 GB) Recommended
Mac Studio M2 Ultra (64 GB+) Ideal for production

See Also

Quant trade-off (MLX lane)

Bits Approx size Use case Recommendation
2-bit ~2.6 GB Aggressive quantization Very low-RAM Macs
3-bit ~3.6 GB Lossy but small Low-RAM Macs
4-bit ~4.2 GB Balanced default Recommended for most Macs
5-bit ~5.0 GB Higher fidelity Quality-sensitive
6-bit ~6.0 GB Approaching FP16 quality High-fidelity
8-bit ~7.6 GB Near-lossless reference Fidelity-critical work

(Current variant β€” 4bit β€” is bolded.)

Variants in this family

(Showing 8 sibling variants under majentik/meralion2-10b-*. The current variant β€” RotorQuant-MLX-4bit β€” is bolded.)

Variant Runtime Approx size Use case
RotorQuant-MLX-4bit mlx-lm ~6.2 GB Apple Silicon balanced

About the RotorQuant / TurboQuant labels

RotorQuant and TurboQuant are this project's release labels, not distinct quantization algorithms β€” for any given tier, both brand repos carry byte-identical weights produced with the standard MLX / llama.cpp quantizers. No brand-specific speedup is claimed or measured.

Downloads last month
50
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Collection including majentik/MERaLiON-2-10B-RotorQuant-MLX-4bit