File size: 8,673 Bytes
ddf8c5b
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
1
2
3
4
5
6
7
8
9
10
11
12
13
14
15
16
17
18
19
20
21
22
23
24
25
26
27
28
29
30
31
32
33
34
35
36
37
38
39
40
41
42
43
44
45
46
47
48
49
50
51
52
53
54
55
56
57
58
59
60
61
62
63
64
65
66
67
68
69
70
71
72
73
74
75
76
77
78
79
80
81
82
83
84
85
86
87
88
89
90
91
92
93
94
95
96
97
98
99
100
101
102
103
104
105
106
107
108
109
110
111
112
113
114
115
116
117
118
119
120
121
122
123
124
125
126
127
128
129
130
131
132
133
134
135
136
137
138
139
140
141
142
143
# K3 Rental Validation β€” Durable Findings & Fixes

Hard-won root causes, patches, and decisions. Transient state (PIDs, SHAs, run progress)
lives in the session log, not here.

## 1. K3 + DSpark speculative decoding crash β†’ FIXED (patch)

**Symptom:** llama-server with `--spec-type draft-dspark -md draft.gguf` loads fine,
reaches READY, but crashes on the FIRST decode:
```
llama-graph.cpp:1376: GGML_ASSERT(t_layer_inp[il] != nullptr && "layer input tensor is null") failed
```

**Root cause:** DSpark/dflash speculative decoding extracts intermediate "layer input"
features from the target (K3) model to feed the draft. It taps specific layers via
`cparams.embeddings_layer_inp[]` (the Lucebox draft GGUF requests `target_layer_ids =
[7, 23, 51, 67, 83]`). At `set_outputs()` time, llama.cpp asserts every requested layer
has a non-null `res->t_layer_inp[il]`. The kimi-k3 architecture graph builder
(`src/models/kimi-k3.cpp`) **never populates `res->t_layer_inp[]`** β€” so the assert fires.
Architectures that DO populate it (and thus support DSpark): `deepseek4.cpp`,
`bailingmoe3.cpp`, `gemma4.cpp`.

**Fix (applied on box `/root/llama.cpp`, backup at `kimi-k3.cpp.bak`):** in the kimi-k3
graph builder layer loop, immediately after `const auto & layer = model.layers[il];`:
```cpp
// expose the raw layer input for speculative draft (DSpark/dflash) feature taps
if ((size_t) il < cparams.embeddings_layer_inp.size() && cparams.embeddings_layer_inp[il]) {
    res->t_layer_inp[il] = inpL;
    cb(res->t_layer_inp[il], "layer_inp", il);
    ggml_build_forward_expand(gf, res->t_layer_inp[il]);
}
```
K3's `inpL` is already the plain per-layer input (no hyper-connection transform needed,
unlike deepseek4's `dsv4_hc_mean`). Draft taps are all at il<93, so no post-loop tail
extraction required. Rebuild: `cmake --build . --config Release -j 56 --target llama-server`.
Result: DSpark decodes on K3 without crash. **Upstream PR candidate.**

## 1b. DSpark draft fails every step "invalid token[1] = -1" β†’ FIXED (GGUF mask token)

**Symptom:** After the t_layer_inp patch, the server loads + the main model decodes, but
the draft fails EVERY step: `init: invalid token[1] = -1` β†’ `decode: failed to initialize
batch` β†’ `llama_decode returned -1` β†’ `draft: llama_decode returned -1`. Main model still
generates (falls back to no speculation) β†’ runs at baseline speed with extra overhead,
zero spec speedup.

**Root cause:** DSpark builds the draft batch with a mask token for the multi-token block:
`common_batch_add(batch, i==0 ? dp.id_last : mask_token_id, ...)` (speculative.cpp:1189).
It resolves the mask id via `llama_vocab_mask(vocab)` which reads
`tokenizer.ggml.mask_token_id`. Our rewritten draft GGUF copied K3's tokenizer keys
(`fix_draft_gguf.py`) but K3 has NO mask token β†’ `llama_vocab_mask()` returns -1 β†’ batch
fed token -1 β†’ embedding rejects it. The draft's TRAINED mask id lives in
`dflash.mask_token_id` = **163824** (separate metadata key, present in both original and
rewritten draft), but llama.cpp never reads it for the mask.

**Fix:** add `tokenizer.ggml.mask_token_id = 163824` (uint32) to the draft GGUF KV section
(`add_mask_token.py`, value taken from `dflash.mask_token_id`). Output
`draft_masked.gguf`; repoint `draft.gguf` symlink at it. Log then shows
`mask_token_id=163824` and zero draft errors. (Arguably also a mainline improvement:
fall back to `dflash.mask_token_id` when the vocab has no mask token.)

## 2. DSpark draft-model VRAM OOM β†’ fixed with `-ngld 0`

The 2.4GB Q8_0 DSpark draft tried to allocate a 4GB KV cache on GPU device 1, which is
already near-full from the K3 trunk 8-way split. Fix: `-ngld 0` (`--gpu-layers-draft 0`)
runs the draft entirely on CPU. Draft is tiny; the main K3 trunk is the bottleneck anyway.

## 3. `-fa on` + large batch OOM on 16GB cards β†’ reduced batch + device-0 share

With `--tensor-split 0.3,1,1,1,1,1,1,1 -b 512 -ub 512`, enabling `-fa on` OOMs the
compute pp buffers. Kitchen-sink config that reaches READY: `--tensor-split
0.2,1,1,1,1,1,1,1 -b 256 -ub 256 -fa on`. (Without `-fa on`, batch 512 + split 0.3 also
OOMs compute pp buffers; the no-FA auto path was the previously-working config.)

## 4. Load-time OOM (RssFile 482GB→cgroup 503GB) → `LLAMA_MMAP_NO_PREFETCH=1`

`src/llama-model.cpp:1663` calls `ml.init_mappings(true, ...)` β†’ MAP_POPULATE +
MADV_WILLNEED eagerly faults all 1.5TB expert pages at load. Patch reads
`LLAMA_MMAP_NO_PREFETCH` env β†’ `init_mappings(!no_prefetch, ...)`. Result: RssFile 124GB
at load, lazy LRU page-cache becomes the hot-expert cache (the architecture the user wants).
This patch is REQUIRED on the home 768GB box too β€” stock llama.cpp cannot load Q4_K_XL.

## 5. HF download throttling β†’ presigned URL + aria2c

`curl -L` on huggingface.co/resolve is throttled to ~7KB/s on many datacenter routes.
Resolve without `-L`, extract the presigned CloudFront URL, download with `aria2c -x16`
β†’ 255-281 MB/s. (`lib/hf_direct.sh`.)

## 6. Bash `GROUPS` is special β€” never use it as a var name

`GROUPS` is a bash builtin array (user's group IDs). Using it for suite group selection
caused a silent no-op. Renamed to `SEL`.

## 7. The trunk is Q8_0, NOT 4-bit β€” and DSpark-on-CPU is a net loss (both fixed by Q4 trunk)

**Discovery (user's instinct was right):** "UD-Q4_K_XL" only 4-bit-quantizes the **routed
experts** (MXFP4). The entire **trunk β€” attention, shared experts (`_shexp`), output head,
token embedding β€” is Q8_0** (8-bit), plus F32 norms. Measured across all 32 shards:
- trunk Q8_0 = 59.6 GB (1116 tensors), norms F32 = 2.6 GB, experts = MXFP4 (rest of 1.4TB).

**Why:** K3 is QAT-trained in MXFP4 β€” the 4-bit experts ARE the reference model (no BF16
original). But the **non-expert path stays higher-precision** (activations MXFP8, non-expert
weights higher precision). So Q8_0 trunk is a legit DOWN-quant from a higher-precision
source, NOT an up-quant of 4-bit. Ref: dreaming.press "Kimi K3's Weights Are Already 4-Bit".
Consequence: re-quantizing the TRUNK Q8_0β†’Q4_K is valid (source was >4-bit). NEVER re-quant
the MXFP4 experts (destroys QAT calibration).

**DSpark-on-CPU measured result:** kitchen sink with `-ngld 0` (draft on CPU) gave
**ks_cold tg 0.224 t/s vs ~0.5 t/s no-spec baseline = ~2x LOSS**. Decode is pinned by CPU
expert execution; the CPU draft forward steals the same 112 threads. DSpark only wins with
the draft ON GPU, but the 58GB Q8_0 trunk fills all 8Γ—16GB cards. β†’ Need a smaller trunk.

**The unlock (both home fit + GPU DSpark):** requant trunk Q8_0β†’Q4_K.
- Trunk 59.6β†’31.6 GB; GPU-resident 62.2β†’34.2 GB. Fits home 2Γ—3090 (48GB) AND frees ~28GB
  on the rental box β†’ DSpark draft can go on GPU.
- **llama-quantize CANNOT do this safely**: `--allow-requantize` forces every non-overridden
  tensor (incl. MXFP4 experts) to the positional type β†’ dequant+requant experts β†’ QAT loss.
  Dry-run confirmed experts became q4_K/q6_K and total size GREW 1438771β†’1602481 MiB.
- **Solution: custom surgical rewriter `requant_trunk.c`** (compiled on box at
  `/root/k3-test/requant_trunk`). Per-shard, split-in=split-out. Byte-copies MXFP4 experts
  + F32 norms unchanged; dequant Q8_0β†’F32 + requant F32β†’Q4_K (ggml
  `dequantize_row_q8_0`/`quantize_row_q4_K_ref`) for trunk tensors only. Falls back to copy
  for tensors whose dims[0] not divisible by 256 (e.g. attn_k_b/ssm_f_b at 128).
  Validated on shard 2: 43 requant + 62 copy, 47.5β†’44.9 GB, no crash.

## Baselines (mainline, batch-1, 64in/64out unless noted)

- Best known (t112, b512/ub512, split 0.3,1..., FA-auto/off, no spec): cold pp ~0.30-0.40
  tg ~0.33-0.51; warm pp 2.5-2.8 tg 0.65-1.79 t/s.
- DSpark-on-CPU kitchen sink: cold pp 0.237 tg 0.224 t/s (~2x WORSE than no-spec).
- t56 regression: warm pp 0.62 tg ~0.49-0.59 β†’ **t112 wins** (SMT siblings help).
- Decode is pinned by CPU expert execution, not cache warm-up (warm tg β‰ˆ cold tg).

## Target architecture (home build)

- 2Γ— 3090 24GB (48GB VRAM) + 768GB RAM + EPYC. Trunk Q8_0 58GB > 48GB β†’ **Q4_K trunk
  requant (34.2GB incl. norms) fits with room for the DSpark draft on GPU.** This is the
  chosen path (see #7). Untestable for exact 2Γ—3090 split on the 8Γ—16GB rental box, but
  the Q4_K trunk + GPU-draft DSpark combo IS testable there (frees ~28GB VRAM).
- Experts: as many as fit in RAM (~150 hottest mlock-pinned, deferred), rest lazily
  faulted/evicted from SSD via kernel page-LRU (`--cpu-moe` + lazy mmap). Q4 only
  (no smaller quant β€” quality). Shared experts (`_shexp`) are GPU-resident, never evicted.
- Spec (DSpark, now working) + expert offload combined = the realistic production case.