Instructions to use vvsotnikov/Qwen3.8-27B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use vvsotnikov/Qwen3.8-27B-MLX-4bit with MLX:
# Make sure mlx-vlm is installed # pip install --upgrade mlx-vlm from mlx_vlm import load, generate from mlx_vlm.prompt_utils import apply_chat_template from mlx_vlm.utils import load_config # Load the model model, processor = load("vvsotnikov/Qwen3.8-27B-MLX-4bit") config = load_config("vvsotnikov/Qwen3.8-27B-MLX-4bit") # Prepare input image = ["http://images.cocodataset.org/val2017/000000039769.jpg"] prompt = "Describe this image." # Apply chat template formatted_prompt = apply_chat_template( processor, config, prompt, num_images=1 ) # Generate output output = generate(model, processor, formatted_prompt, image) print(output) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use vvsotnikov/Qwen3.8-27B-MLX-4bit with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "vvsotnikov/Qwen3.8-27B-MLX-4bit"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "vvsotnikov/Qwen3.8-27B-MLX-4bit" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use vvsotnikov/Qwen3.8-27B-MLX-4bit with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "vvsotnikov/Qwen3.8-27B-MLX-4bit"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "vvsotnikov/Qwen3.8-27B-MLX-4bit" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Hermes Agent
How to use vvsotnikov/Qwen3.8-27B-MLX-4bit with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "vvsotnikov/Qwen3.8-27B-MLX-4bit"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default vvsotnikov/Qwen3.8-27B-MLX-4bit
Run Hermes
hermes
- Atomic Chat
Qwen3.8-27B-MLX-4bit
This model was converted to MLX format from
Qwen/Qwen3.8-27B using mlx-vlm version
0.6.8, and quantized to 4 bits, so it runs on Apple
silicon through mlx-vlm. Refer to the
original model card for more details
on the model itself.
Pair it with
Qwen3.8-27B-MTP-MLX-4bit
to get speculative decoding, because Qwen carries a Multi-Token Prediction
head that the MLX conversion moves into a separate repository.
Use with mlx-vlm
pip install -U mlx-vlm
mlx_vlm generate \
--model vvsotnikov/Qwen3.8-27B-MLX-4bit \
--draft-model vvsotnikov/Qwen3.8-27B-MTP-MLX-4bit \
--prompt "Write a quicksort in Python." \
--max-tokens 256 --temperature 0.6 --enable-thinking
Drop --draft-model to decode without the speculator.
How this was produced
mlx_vlm convert --hf-path Qwen/Qwen3.8-27B \
--mlx-path Qwen3.8-27B-MLX-4bit -q --q-bits 4 --q-group-size 64
The quantization is uniform affine at 4 bits with group size
64, and it covers 498 modules
of the language model including embed_tokens, lm_head and the
linear_attn projections. The vision tower stays dense in bfloat16, since
the converter skips multimodal modules by default.
Verification
Every artifact in this release passed an automated check before it was published, and nothing became public until its remote checksums matched the local build.
| Check | Result |
|---|---|
| Quantization | {bits: 4, group_size: 64, mode: affine} |
| Quantized modules | 498 of 2180 tensors |
| Vision tower | dense, 333 tensors |
| MTP tensors | 0, moved to the drafter repository |
| Acceptance with the 4-bit drafter | 93.5% of drafted tokens accepted, 2.86 accepted tokens/round, over 70 rounds, 15.982 tok/s |
License and attribution
The weights derive from Qwen/Qwen3.8-27B
under Apache 2.0, so the original license and its terms carry over.
Read the license itself before you use this model,
and refer to the upstream model card for the model's capabilities and limits.
- Downloads last month
- 447
4-bit
Model tree for vvsotnikov/Qwen3.8-27B-MLX-4bit
Base model
Qwen/Qwen3.8-27B