burmese-coder-4b-mlx / README_mlx.md
WYNN747's picture
Upload folder using huggingface_hub
c6b58f8 verified
|
Raw
History Blame Contribute Delete
1.5 kB

Burmese Coder 4B - MLX (4-bit Quantized)

This directory contains the Burmese Coder 4B model converted to the native MLX format, optimized for Apple Silicon (M1/M2/M3/M4).

🚀 Performance

  • Format: MLX (4-bit quantization)
  • Size: ~2.2 GB
  • Inference Speed: ~60 tokens/sec (on standard Apple Silicon)
  • Memory Usage: ~2.3 GB Peak VRAM

🛠️ Prerequisites

Ensure you have mlx-lm installed in your Python environment (Python 3.10+ required):

pip install mlx-lm

💬 Usage (CLI Inference)

You can run the model directly from the command line:

python -m mlx_lm.generate --model burmese_coder_4b_mlx --prompt "Python မှာ list တစ်ခုကို ဘယ်လို sort လုပ်ရသလဲ?" --max-tokens 300

🐍 Python API Usage

from mlx_lm import load, generate

model, tokenizer = load("burmese_coder_4b_mlx")

prompt = "Python မှာ list တစ်ခုကို ဘယ်လို sort လုပ်ရသလဲ?"
messages = [{"role": "user", "content": prompt}]
prompt = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)

response = generate(model, tokenizer, prompt=prompt, verbose=True)
print(response)

📝 Conversion Details

  • Base Model: unsloth/gemma-3-4b-it
  • Adapter: WYNN747/burmese-coder-4b (LoRA)
  • Quantization: 4-bit (--q-bits 4)
  • Fusing: Merged using mlx-lm fuse tool before quantization for maximum integrity.