Instructions to use Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality"
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- MLX LM
How to use Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality
Run Hermes
hermes
- Atomic Chat
Qwen3.8-27B MTPLX Optimized Quality
⏳ Placeholder — weights pending
The official Qwen/Qwen3.8-27B release is scheduled for 2026-08-14 15:00 UTC. This build starts the moment the official weights are readable. Like/watch this repo to get it the moment it's live. License will follow the upstream Qwen3.8-27B license.
The highest-fidelity MTPLX build of Qwen3.8-27B for Apple Silicon — for when what the model says matters more than how fast it says it.
Optimized Quality is the build for long agent sessions and the hardest tasks: staying closest to the original model's behavior while still decoding multiple tokens per step through the model's native multi-token-prediction head — preserved on load, verified with exact rejection sampling, so the output distribution matches plain decoding at real sampling settings. No greedy shortcut.
How this build is composed gets decided the way every MTPLX release is: measured on the real weights, on real hardware, against the alternatives — then shipped. Recipe details land here with the artifact, not before.
Quickstart (once weights land)
brew install youssofal/mtplx/mtplx
mtplx pull Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality
mtplx run "hello" --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality
mtplx serve --model Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality --port 8000
mtplx serve exposes OpenAI- and Anthropic-compatible endpoints, so the
model works in anything that speaks either API.
The numbers
Posted here after the build, from max-fan verified runs on real hardware, with the workload named — the same discipline as every MTPLX release. No number on this card will predate the weights.
The MTPLX family for Qwen3.8-27B
| Build | Best for |
|---|---|
| Bare Speed | fastest short-context chat |
| Optimized Speed | fast coding + agents |
| Optimized Quality (this repo) | highest fidelity, long agent sessions |
- Downloads last month
- -
8-bit
Model tree for Youssofal/Qwen3.8-27B-MTPLX-Optimized-Quality
Base model
Qwen/Qwen3.8-27B