Qwen3.6-27B-DFlash2

This is a DFlash2 draft model intended for speculative decoding alongside its separately distributed, unmodified target model. It is not a standalone replacement for the target model.

Model details

  • Architecture: DFlash2DraftModel
  • Precision: bfloat16
  • Context length: 262,144 tokens
  • Draft blocks: 8 tokens per block
  • Draft selector: top 16 candidates, rank 256
  • Target-layer taps: 5, 19, 33, 47, and 61

The model uses the supplied Qwen tokenizer and a sliding attention window of 2,048 tokens.

Training

The drafter was fine-tuned for one epoch on a domain-specific corpus of image-prompt-to-Three.js-module generations, using a learning rate of 6e-5. Only the drafter weights were changed. The tokenizer, DFlash2 configuration, and target-layer taps remain unchanged.

Usage

Load this repository with a runtime that supports DFlash2DraftModel and pair it with the corresponding target model for speculative decoding. Standard transformers loading alone may not provide the required DFlash2 runtime.

License and attribution

This repository is a derivative of Qwen/Qwen3.8-27B and is distributed under the Apache License 2.0. See NOTICE for the required change notice and LICENSE for the full license text.

Downloads last month
13
Safetensors
Model size
2B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Fluxmire/Qwen3.6-27B-DFlash2

Base model

Qwen/Qwen3.8-27B
Finetuned
(310)
this model