Image-Text-to-Text
MLX
Safetensors
mage_vl
apple-silicon
vision-language
video
streaming
conversational
custom_code
4-bit precision
Instructions to use sr29/Mage-VL-mlx-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use sr29/Mage-VL-mlx-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("sr29/Mage-VL-mlx-4bit") config = load_config("sr29/Mage-VL-mlx-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Update README.md
Browse files
README.md
CHANGED
|
@@ -50,7 +50,7 @@ python scripts/generate.py --mlx mage-vl-mlx --tokenizer-src mage-vl-mlx \
|
|
| 50 |
|
| 51 |
4-bit is recommended for 16GB; 8-bit gives richer output if you have RAM.
|
| 52 |
|
| 53 |
-
## Validation
|
| 54 |
|
| 55 |
- **Image preprocessing**: bit-exact vs the HF `Qwen2VLImageProcessor` (max_abs_diff 0.0).
|
| 56 |
- **Vision tower**: numerically matches the reference weights end-to-end
|
|
|
|
| 50 |
|
| 51 |
4-bit is recommended for 16GB; 8-bit gives richer output if you have RAM.
|
| 52 |
|
| 53 |
+
## Validation
|
| 54 |
|
| 55 |
- **Image preprocessing**: bit-exact vs the HF `Qwen2VLImageProcessor` (max_abs_diff 0.0).
|
| 56 |
- **Vision tower**: numerically matches the reference weights end-to-end
|