File size: 7,115 Bytes
fd0ee83 01b7820 7a8fcc0 01b7820 7a8fcc0 01b7820 7a8fcc0 0b756c0 7a8fcc0 0b756c0 7a8fcc0 0b756c0 7a8fcc0 0b756c0 7a8fcc0 0b756c0 7a8fcc0 0b756c0 7a8fcc0 0b756c0 7a8fcc0 0b756c0 7a8fcc0 0b756c0 7a8fcc0 01b7820 7a8fcc0 01b7820 7a8fcc0 01b7820 7a8fcc0 0b756c0 7a8fcc0 01b7820 7a8fcc0 01b7820 7a8fcc0 01b7820 7a8fcc0 01b7820 7a8fcc0 | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 144 145 146 147 | ---
license: apache-2.0
base_model:
- Qwen/Qwen3.5-9B
- nightmedia/Qwen3.5-9B-Holodeck-Lounge
- Jackrong/Qwopus3.5-9B-Coder
tags:
- multi-token-prediction
- speculative-decoding
- qwen3.5
- mtp
- merge
---
# Qwen3.5-9B-Holodeck-Lounge (MTP-restored)
> **TL;DR** A working **native Multi-Token Prediction (MTP) head** on top of the
> Holodeck-Lounge merge. The upstream merge lineage shipped **without a
> functional MTP head** (the `mtp.fc.weight` tensor is absent after merge), so
> we restored it by sourcing only that **single missing tensor** β
> `mtp.fc.weight` β from
> [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder).
> The other **14 of the MTP head's 15 layers** (the `mtp.layers.*` blocks) were
> already present in the merge. We also patched in the
> [froggeric/Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
> fixed jinja template (**v22.1**), which improves reasoning
> structure and MTP acceptance in our evals. This model is the reference donor for our
> [head fine-tune experiments](https://huggingface.co/Wiself/Holodeck-Lounge-Regressed-MTP).
## Why this model exists
- [nightmedia/Qwen3.5-9B-Holodeck-Lounge](https://huggingface.co/nightmedia/Qwen3.5-9B-Holodeck-Lounge)
is a fine creative-fiction merge. Excellent writing model β but its merge
lineage (13 base models, see the upstream card for the full recipe) does not
carry a working MTP head: `mtp.fc.weight` is missing, so native speculative
decoding is impossible from the source.
- **What we did:** the merge already carried **14 of the MTP head's 15 layers**
(the `mtp.layers.*` transformer blocks); only the input-fusion projection
`mtp.fc.weight` was absent. We sourced that **single tensor** from
**Jackrong/Qwopus3.5-9B-Coder**, transplanted it into the Holodeck-Lounge
backbone, and verified it loads and runs end-to-end in llama.cpp
(`draft-mtp`) and vLLM (`qwen3_5` MTP method). Backbone weights untouched.
- **Chat template:** baked in the fixed jinja template v22.1 from
[froggeric/Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
(Qwen 3.5/3.6/3.8), improving reasoning
structure, organisation, coherence, and MTP acceptance rate in our evals.
## Native MTP head
- Single-layer, DeepSeek-style MTP head (as in the
[DeepSeek-V3 Technical Report](https://arxiv.org/abs/2412.19437)),
**3 draft tokens**, fixed adjacent context.
- The 14 `mtp.layers.*` blocks come from the Holodeck-Lounge merge; only the
`mtp.fc.weight` fusion projection was sourced from
Jackrong/Qwopus3.5-9B-Coder (a model whose checkpoint carries the same
native Qwen3.5 MTP architecture).
- The donor projection plus the 14 host layers reach the acceptance range
below β higher than any of our head-only fine-tunes
([related experiment β](https://huggingface.co/Wiself/Holodeck-Lounge-Regressed-MTP)).
## Evaluation
| Metric | Value |
|---|---|
| llama.cpp greedy / target-only acceptance | **59.0%** |
| vLLM rejection-sampling acceptance | **58.55%** |
| RS pos0 / pos1 / pos2 | 77.2% / 57.0% / 41.4% |
| Mean accepted length | 2.756 |
Eval: 30 prompts, temperature 0.7, max 256 tokens, concurrency 8. llama.cpp
numbers are greedy target-only acceptance (reasoning off); vLLM RS acceptance
from `vllm:spec_decode_num_accepted_tokens / _num_draft_tokens`. With
reasoning enabled per the Usage section, measured acceptance reaches the
**60β69%** range (n=30, same eval set).
> **Eval caveat:** n=30 is a small sample β treat the 3-decimal precision as
> indicative; confidence intervals are wide.
## Real-world usage
With the recommended settings β `draft-mtp` in llama.cpp, `qwen3_5` MTP in
vLLM β the restored head has consistently delivered a **60β69% draft-token
acceptance rate** (65+% typical) in our testing. Real-world mileage will vary with workload
and hardware; treat 65% as a strong baseline, not a guarantee. The point of
this model isn't raw speed, but **option**: MTP when it pays off, plain greedy
decoding when it doesn't.
Where the win matters most:
- Bandwidth-constrained serving, where accepted draft tokens cut
tokens-per-request at the wire.
- High-concurrency, latency-sensitive inference, where saved decode steps
shrink queue and p95.
- Settings where every accepted draft token compounds β long generations,
reasoning-heavy prompts, tight SLOs.
The clean part: the host checkpoint already carried **14 of the MTP head's 15
layers**, so this model ships at essentially the same size as its host β it was
always that big. The single-tensor transplant just completes the missing
`mtp.fc.weight` fusion layer, so you get the choice in one model, at no real
size cost.
Thanks to nightmedia for the Holodeck-Lounge backbone (which already carried 14
of the MTP head's 15 layers), to Jackrong/Qwopus3.5-9B-Coder for the
`mtp.fc.weight` tensor that completed the restore, and to froggeric for the
Qwen-Fixed-Chat-Templates v22.1 baked in here.
## Usage (recommended)
For best performance, use the baked-in froggeric template with reasoning
enabled and preserved:
```bash
llama-server -m model.gguf \
--jinja --chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning on --reasoning-preserve \
--spec-type draft-mtp --spec-draft-n-max 3
```
- **llama.cpp MTP:** `--spec-type draft-mtp --spec-draft-n-max 3`
- **vLLM MTP:** `--speculative-config '{"method":"mtp","num_speculative_tokens":3}'`
## Provenance & credits
| Component | Source |
|---|---|
| Base architecture | [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) (Qwen team) |
| Backbone merge | [nightmedia/Qwen3.5-9B-Holodeck-Lounge](https://huggingface.co/nightmedia/Qwen3.5-9B-Holodeck-Lounge) β 13-model creative-writing merge (DavidAU heretic family, armand0e, microsoft/Fara1.5-9B, Jackrong, et al.; see upstream card) |
| MTP head β 14 `mtp.layers.*` blocks + `mtp.norm`/`mtp.pre_fc_norm_*` | Present in [nightmedia/Qwen3.5-9B-Holodeck-Lounge](https://huggingface.co/nightmedia/Qwen3.5-9B-Holodeck-Lounge) merge |
| MTP head β `mtp.fc.weight` (input-fusion projection, the missing 15th layer) | [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) |
| Chat template (v22.1) | [froggeric/Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates) |
License: **apache-2.0**. License chain (all components apache-2.0, verified):
Qwen/Qwen3.5-9B β nightmedia merge β Jackrong MTP head β froggeric template.
## Limitations
- Acceptance trails dedicated trained draft heads such as
[EAGLE-3](https://arxiv.org/abs/2503.01840) (β61% on Qwen3.5-9B).
- Deepest draft positions (pos2) accept at lower rates than pos0/pos1.
- The transplanted head is inference-accelerator only; it does not affect
backbone text quality.
- The head is only directly usable as a built-in MTP head for this
architecture (same hidden size + vocab); as a standalone GGUF it can also be
used as a sidecar draft (`--model-draft mtp-*.gguf`).
**AI was used to draft this report.** |