Qwen3.8-27B DFlash2 โ€” MLXFast affine-4 artifact

This repository contains a proposal-only, MLX affine-4 quantization of incoai/Qwen3.8-27B-DFlash2 for the Yukon MLXFast Qwen3.8 challenge. It is not a standalone language model. A Qwen3.8-27B target model must verify every proposed token.

The original DFlash2 checkpoint and its z-lab/Qwen3.8-27B-DFlash2 mirror contain the same model.safetensors blob. This derived artifact was made from that exact public blob at the following immutable revisions:

  • incoai/Qwen3.8-27B-DFlash2@adde41d8fde3a75dc905a7df0bd5088d2a44b5a1
  • z-lab/Qwen3.8-27B-DFlash2@ac04198556d7e8867853cbc356807b969f311b05

Quantization

  • Format: MLX affine quantization
  • Linear weights: 4 bits, group size 64
  • Non-linear parameters and selector codebooks: BF16
  • Tensor count: 175
  • Quantized linear modules: 47
  • model.safetensors bytes: 1,265,644,781
  • model.safetensors SHA-256: 95cd528f14ced30c4cb4933c2ae233f822c20cc6c38f7502adccef359c0141ce

The converted tensor tree was strict-loaded into a faithful MLX DFlash2 implementation before publication. The Yukon submission independently pins this file by immutable Hub revision, exact byte count, and the challenge's documented single-file tree digest.

Architecture retained

The artifact preserves the released five-layer DFlash2 drafter, including:

  • target features from layers 5, 19, 33, 47, and 61;
  • parallel block drafting with the target embedding and output head;
  • 2-tap dynamic grouped convolutions before and after attention and MLP;
  • a top-16, rank-256 adjacent-candidate path selector;
  • sliding-window grouped-query attention.

Every output token remains target-authoritative: the draft is only a proposal, and the challenge's trusted parent verifies the full target trajectory.

Upstream

This derived artifact is published under the upstream Apache-2.0 license.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
BF16
ยท
U32
ยท
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for ProCreations/Qwen3.8-27B-DFlash2-MLXFast-Q4

Base model

Qwen/Qwen3.8-27B
Finetuned
(142)
this model