Image-Text-to-Text
MLX
Safetensors
mage_vl
apple-silicon
vision-language
video
streaming
conversational
custom_code
4-bit precision
Instructions to use sr29/Mage-VL-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use sr29/Mage-VL-mlx-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("sr29/Mage-VL-mlx-4bit") config = load_config("sr29/Mage-VL-mlx-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
| license: apache-2.0 | |
| base_model: microsoft/Mage-VL | |
| tags: | |
| - mlx | |
| - apple-silicon | |
| - vision-language | |
| - image-text-to-text | |
| - video | |
| - streaming | |
| pipeline_tag: image-text-to-text | |
| library_name: mlx | |
| # Mage-VL 4-bit (MLX, Apple Silicon) | |
| An **MLX** (4-bit) conversion of [microsoft/Mage-VL](https://huggingface.co/microsoft/Mage-VL) — | |
| a 5B image/video vision-language model (Qwen3-4B text backbone + a from-scratch | |
| **Mage-ViT** "Codec-ViT" vision encoder) — that runs on Apple Silicon. | |
| Port code, converter, and validators: **https://github.com/rsravanreddy/Mage-VL-MLX** | |
| - Weights: `3.1 GB` (4-bit, group size 64) | |
| - Runs image, video (frame-sampling), and the **streaming event gate** locally on a Mac. | |
| ## Usage | |
| ```bash | |
| # 1. install the MLX stack + register the mage_vl plugin | |
| pip install mlx mlx-lm mlx-vlm numpy pillow tokenizers jinja2 av | |
| git clone https://github.com/rsravanreddy/Mage-VL-MLX && cd Mage-VL-MLX | |
| ln -s "$PWD/mage_vl" "$(python -c 'import mlx_vlm,os;print(os.path.dirname(mlx_vlm.__file__))')/models/mage_vl" | |
| # 2. download these weights | |
| hf download sr29/Mage-VL-mlx-4bit --local-dir mage-vl-mlx | |
| # 3. run (image or video) | |
| python scripts/generate.py --mlx mage-vl-mlx --tokenizer-src mage-vl-mlx \ | |
| --image path/to/image.jpg --prompt "Describe this image." | |
| python scripts/generate.py --mlx mage-vl-mlx --tokenizer-src mage-vl-mlx \ | |
| --video path/to/video.mp4 --num-frames 8 --prompt "What is happening?" | |
| ``` | |
| ## Performance (Apple M4, 16GB) | |
| | model | weights | image decode | image peak RAM | | |
| |-------|--------:|-------------:|---------------:| | |
| | 4-bit | 3.1 GB | 30.6 tok/s | 4.65 GB | | |
| | 8-bit | 5.0 GB | 19.1 tok/s | 6.55 GB | | |
| 4-bit is recommended for 16GB; 8-bit gives richer output if you have RAM. | |
| ## Validation | |
| - **Image preprocessing**: bit-exact vs the HF `Qwen2VLImageProcessor` (max_abs_diff 0.0). | |
| - **Vision tower**: numerically matches the reference weights end-to-end | |
| (`max_abs_diff 3.0e-4`, fp32, full 24 layers). | |
| - **Streaming Mamba mixer**: matches a canonical selective-scan reference (`4.3e-7`). | |
| - **Generation**: qualitatively correct on image + video. | |
| - **Not yet done**: full end-to-end logit parity vs the HF model (memory-gated); | |
| the codec (token-reduction) video backend needs the external codec engine. | |
| ## Streaming | |
| Mage-VL's proactive **event gate** (`streammind_gate`) is ported | |
| (`mage_vl/streaming.py`). See `scripts/stream.py` for a per-frame silent/speak | |
| timeline. Note: the gate weights (`streammind_gate.safetensors`) are separate and | |
| downloaded from the upstream Mage-VL repo. | |
| ## License & attribution | |
| Apache-2.0. Derivative of [microsoft/Mage-VL](https://huggingface.co/microsoft/Mage-VL) | |
| (Apache-2.0); reuses the Qwen3 language model from | |
| [mlx-vlm](https://github.com/Blaizzy/mlx-vlm). Weights converted, not retrained. | |