Qwen3.8-Flash-Next-4bit-paged

Expert-paged build of Vontra/Qwen3.8-Flash-Next-MLX-4bit. The weights that are read a fraction at a time live in their own containers, so a machine loads what it needs rather than all of it.

file size holds
model.safetensors 3.80 GiB resident weights
experts.bin 70.31 GiB routed experts
ple-q4.rows 29.80 GiB n-gram table
mtp/ 1.52 GiB draft head, off by default

Total 105.46 GiB. Of that, 103.94 GiB is the source build, whose bytes moved into

These containers are not a format mlx-lm reads. The model runs on gbx_lm, a single signed binary for Apple Silicon; there is nothing to pip install. containers rather than being copied, and 1.52 GiB is the draft head, which no published build of this model carries.

Requirements

macOS 15.0 or later
chip Apple Silicon (arm64). There is no Intel build.
Python none -- the binary carries what it needs

Memory is not a fixed figure for a paged build, and that is the point of one: it fills what fits and streams the rest from disk. On a 512 GB Mac Studio with room to spare this model settles at about 76 GB resident. A smaller machine holds less and reads more from disk -- slower, but it runs.

How much slower depends on how far the machine is from holding the experts, and on how fast its disk is. Each token routes to a few experts; the ones already in memory cost nothing to reach, and the ones that are not have to be read before that token can finish. A machine holding most of them waits rarely, one holding few waits often. We have not measured this across machine sizes and will not guess a figure: what we can say is that the model answers either way, and that the wait is the SSD's, not the model's.

Install

# upgrading? clear the previous version's unpack directory first
rm -rf ~/.libra/cache/onefile/gbx_lm

curl -fL -o gbx_lm-darwin-arm64.tar.gz 'https://github.com/GreenBitAI/gbx-lm/releases/latest/download/gbx_lm-darwin-arm64.tar.gz' \
  && tar -xzf gbx_lm-darwin-arm64.tar.gz gbx_lm \
  && mkdir -p "$HOME/.local/bin" \
  && mv gbx_lm "$HOME/.local/bin/gbx_lm" \
  && chmod +x "$HOME/.local/bin/gbx_lm"

gbx_lm -h

The build is signed with a Developer ID and notarised, so macOS runs it without the usual detour for a downloaded binary.

command not found -- $HOME/.local/bin is not on your PATH:

echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.zshrc && source ~/.zshrc     # zsh
echo 'export PATH="$HOME/.local/bin:$PATH"' >> ~/.bash_profile && source ~/.bash_profile   # bash

Killed: 9 -- a previous version's files are still in the unpack directory, and macOS refuses to mix two builds. Run the rm -rf line above, then try again.

Run

gbx_lm --model GreenBitAI/Qwen3.8-Flash-Next-4bit-paged

That serves an OpenAI-compatible API on port 11688, which is its default. The weights download on first use into ~/.libra/cache/models; set HF_HOME to put them elsewhere, and HF_TOKEN if you meet the Hub's rate limits for anonymous downloads.

curl http://127.0.0.1:11688/v1/chat/completions \
  -H 'Content-Type: application/json' \
  -d '{"model":"GreenBitAI/Qwen3.8-Flash-Next-4bit-paged","messages":[{"role":"user","content":"Hello"}]}'

Where the weights fit they are filled from experts.bin and the model runs the stock path at stock speed; where they do not, they stream from disk. Reading the machine decides that, not a flag.

To override that: GBX_PAGING=off holds the experts resident, GBX_PLE=off holds the n-gram table resident.

Checked at build time, while the source checkpoint was still there to compare against:

  • PASS bit-identical logits — exact on 5 prompt(s) to 160 tokens; 1 longer differ by at most 6.8, same greedy token throughout
  • PASS layer-wise vs resident — 48 layers x 2 draws exact, 88.89 GiB peak for this gate

Quantization, tokenizer, chat template and licence are unchanged from Vontra/Qwen3.8-Flash-Next-MLX-4bit.

Coding agents

The server speaks three wire protocols on the same port, so the tools that expect a hosted API can be pointed at this one:

path for
/v1/chat/completions anything written against the OpenAI API
/v1/responses Codex
/v1/messages Claude Code

Codex -- a provider in ~/.codex/config.toml:

[model_providers.gbx]
name = "gbx-lm"
base_url = "http://127.0.0.1:11688/v1"
wire_api = "responses"

and a profile in ~/.codex/gbx.config.toml:

model_provider = "gbx"
model = "GreenBitAI/Qwen3.8-Flash-Next-4bit-paged"
model_context_window = 262144

Claude Code -- ~/.claude/gbx.settings.json:

{
  "env": {
    "ANTHROPIC_BASE_URL": "http://127.0.0.1:11688",
    "ANTHROPIC_AUTH_TOKEN": "local",
    "ANTHROPIC_MODEL": "GreenBitAI/Qwen3.8-Flash-Next-4bit-paged",
    "ANTHROPIC_DEFAULT_HAIKU_MODEL": "GreenBitAI/Qwen3.8-Flash-Next-4bit-paged"
  }
}

Both clients ask for a small model for their own background work, so every name in the settings has to be one this server is serving.

The draft head

The mtp/ folder carries the model's own multi-token prediction head, so speculative decoding works from this repository alone. It is off unless asked for:

GBX_QWEN4_MTP=on gbx_lm --model GreenBitAI/Qwen3.8-Flash-Next-4bit-paged

Up to 2.44x. Measured 2026-09-18 on a 512 GB Mac Studio (M3 Ultra), 128 tokens, greedy, decode timed from the first token; the median of three runs, which agreed to within 1.5%:

context head off head on speedup acceptance
4,096 24.4 tok/s 59.6 2.44x 0.81
16,384 23.9 49.6 2.08x 0.68
30,000 23.5 51.8 2.20x 0.73

Every token the head proposes is checked by the model itself, so the reply is the model's own either way; the head only saves passes over the weights.

Downloads last month
1,053
Safetensors
Model size
5B params
Tensor type
U32
·
BF16
·
I64
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for GreenBitAI/Qwen3.8-Flash-Next-4bit-paged

Quantized
(295)
this model

Collection including GreenBitAI/Qwen3.8-Flash-Next-4bit-paged