Holodeck-Lounge-MTP / README.md
Wiself's picture
Upload README.md
0b756c0 verified
|
Raw
History Blame Contribute Delete
7.12 kB
---
license: apache-2.0
base_model:
- Qwen/Qwen3.5-9B
- nightmedia/Qwen3.5-9B-Holodeck-Lounge
- Jackrong/Qwopus3.5-9B-Coder
tags:
- multi-token-prediction
- speculative-decoding
- qwen3.5
- mtp
- merge
---
# Qwen3.5-9B-Holodeck-Lounge (MTP-restored)
> **TL;DR** A working **native Multi-Token Prediction (MTP) head** on top of the
> Holodeck-Lounge merge. The upstream merge lineage shipped **without a
> functional MTP head** (the `mtp.fc.weight` tensor is absent after merge), so
> we restored it by sourcing only that **single missing tensor**
> `mtp.fc.weight` — from
> [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder).
> The other **14 of the MTP head's 15 layers** (the `mtp.layers.*` blocks) were
> already present in the merge. We also patched in the
> [froggeric/Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
> fixed jinja template (**v22.1**), which improves reasoning
> structure and MTP acceptance in our evals. This model is the reference donor for our
> [head fine-tune experiments](https://huggingface.co/Wiself/Holodeck-Lounge-Regressed-MTP).
## Why this model exists
- [nightmedia/Qwen3.5-9B-Holodeck-Lounge](https://huggingface.co/nightmedia/Qwen3.5-9B-Holodeck-Lounge)
is a fine creative-fiction merge. Excellent writing model — but its merge
lineage (13 base models, see the upstream card for the full recipe) does not
carry a working MTP head: `mtp.fc.weight` is missing, so native speculative
decoding is impossible from the source.
- **What we did:** the merge already carried **14 of the MTP head's 15 layers**
(the `mtp.layers.*` transformer blocks); only the input-fusion projection
`mtp.fc.weight` was absent. We sourced that **single tensor** from
**Jackrong/Qwopus3.5-9B-Coder**, transplanted it into the Holodeck-Lounge
backbone, and verified it loads and runs end-to-end in llama.cpp
(`draft-mtp`) and vLLM (`qwen3_5` MTP method). Backbone weights untouched.
- **Chat template:** baked in the fixed jinja template v22.1 from
[froggeric/Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates)
(Qwen 3.5/3.6/3.8), improving reasoning
structure, organisation, coherence, and MTP acceptance rate in our evals.
## Native MTP head
- Single-layer, DeepSeek-style MTP head (as in the
[DeepSeek-V3 Technical Report](https://arxiv.org/abs/2412.19437)),
**3 draft tokens**, fixed adjacent context.
- The 14 `mtp.layers.*` blocks come from the Holodeck-Lounge merge; only the
`mtp.fc.weight` fusion projection was sourced from
Jackrong/Qwopus3.5-9B-Coder (a model whose checkpoint carries the same
native Qwen3.5 MTP architecture).
- The donor projection plus the 14 host layers reach the acceptance range
below — higher than any of our head-only fine-tunes
([related experiment →](https://huggingface.co/Wiself/Holodeck-Lounge-Regressed-MTP)).
## Evaluation
| Metric | Value |
|---|---|
| llama.cpp greedy / target-only acceptance | **59.0%** |
| vLLM rejection-sampling acceptance | **58.55%** |
| RS pos0 / pos1 / pos2 | 77.2% / 57.0% / 41.4% |
| Mean accepted length | 2.756 |
Eval: 30 prompts, temperature 0.7, max 256 tokens, concurrency 8. llama.cpp
numbers are greedy target-only acceptance (reasoning off); vLLM RS acceptance
from `vllm:spec_decode_num_accepted_tokens / _num_draft_tokens`. With
reasoning enabled per the Usage section, measured acceptance reaches the
**60–69%** range (n=30, same eval set).
> **Eval caveat:** n=30 is a small sample — treat the 3-decimal precision as
> indicative; confidence intervals are wide.
## Real-world usage
With the recommended settings — `draft-mtp` in llama.cpp, `qwen3_5` MTP in
vLLM — the restored head has consistently delivered a **60–69% draft-token
acceptance rate** (65+% typical) in our testing. Real-world mileage will vary with workload
and hardware; treat 65% as a strong baseline, not a guarantee. The point of
this model isn't raw speed, but **option**: MTP when it pays off, plain greedy
decoding when it doesn't.
Where the win matters most:
- Bandwidth-constrained serving, where accepted draft tokens cut
tokens-per-request at the wire.
- High-concurrency, latency-sensitive inference, where saved decode steps
shrink queue and p95.
- Settings where every accepted draft token compounds — long generations,
reasoning-heavy prompts, tight SLOs.
The clean part: the host checkpoint already carried **14 of the MTP head's 15
layers**, so this model ships at essentially the same size as its host — it was
always that big. The single-tensor transplant just completes the missing
`mtp.fc.weight` fusion layer, so you get the choice in one model, at no real
size cost.
Thanks to nightmedia for the Holodeck-Lounge backbone (which already carried 14
of the MTP head's 15 layers), to Jackrong/Qwopus3.5-9B-Coder for the
`mtp.fc.weight` tensor that completed the restore, and to froggeric for the
Qwen-Fixed-Chat-Templates v22.1 baked in here.
## Usage (recommended)
For best performance, use the baked-in froggeric template with reasoning
enabled and preserved:
```bash
llama-server -m model.gguf \
--jinja --chat-template-file chat_template.jinja \
--reasoning-format deepseek --reasoning on --reasoning-preserve \
--spec-type draft-mtp --spec-draft-n-max 3
```
- **llama.cpp MTP:** `--spec-type draft-mtp --spec-draft-n-max 3`
- **vLLM MTP:** `--speculative-config '{"method":"mtp","num_speculative_tokens":3}'`
## Provenance & credits
| Component | Source |
|---|---|
| Base architecture | [Qwen/Qwen3.5-9B](https://huggingface.co/Qwen/Qwen3.5-9B) (Qwen team) |
| Backbone merge | [nightmedia/Qwen3.5-9B-Holodeck-Lounge](https://huggingface.co/nightmedia/Qwen3.5-9B-Holodeck-Lounge) — 13-model creative-writing merge (DavidAU heretic family, armand0e, microsoft/Fara1.5-9B, Jackrong, et al.; see upstream card) |
| MTP head — 14 `mtp.layers.*` blocks + `mtp.norm`/`mtp.pre_fc_norm_*` | Present in [nightmedia/Qwen3.5-9B-Holodeck-Lounge](https://huggingface.co/nightmedia/Qwen3.5-9B-Holodeck-Lounge) merge |
| MTP head — `mtp.fc.weight` (input-fusion projection, the missing 15th layer) | [Jackrong/Qwopus3.5-9B-Coder](https://huggingface.co/Jackrong/Qwopus3.5-9B-Coder) |
| Chat template (v22.1) | [froggeric/Qwen-Fixed-Chat-Templates](https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates) |
License: **apache-2.0**. License chain (all components apache-2.0, verified):
Qwen/Qwen3.5-9B → nightmedia merge → Jackrong MTP head → froggeric template.
## Limitations
- Acceptance trails dedicated trained draft heads such as
[EAGLE-3](https://arxiv.org/abs/2503.01840) (≈61% on Qwen3.5-9B).
- Deepest draft positions (pos2) accept at lower rates than pos0/pos1.
- The transplanted head is inference-accelerator only; it does not affect
backbone text quality.
- The head is only directly usable as a built-in MTP head for this
architecture (same hidden size + vocab); as a standalone GGUF it can also be
used as a sidecar draft (`--model-draft mtp-*.gguf`).
**AI was used to draft this report.**