K3-Stuff / scripts /FINDINGS.md
TessaCoil's picture
Upload folder using huggingface_hub
ddf8c5b verified
|
Raw
History Blame Contribute Delete
8.67 kB

K3 Rental Validation β€” Durable Findings & Fixes

Hard-won root causes, patches, and decisions. Transient state (PIDs, SHAs, run progress) lives in the session log, not here.

1. K3 + DSpark speculative decoding crash β†’ FIXED (patch)

Symptom: llama-server with --spec-type draft-dspark -md draft.gguf loads fine, reaches READY, but crashes on the FIRST decode:

llama-graph.cpp:1376: GGML_ASSERT(t_layer_inp[il] != nullptr && "layer input tensor is null") failed

Root cause: DSpark/dflash speculative decoding extracts intermediate "layer input" features from the target (K3) model to feed the draft. It taps specific layers via cparams.embeddings_layer_inp[] (the Lucebox draft GGUF requests target_layer_ids = [7, 23, 51, 67, 83]). At set_outputs() time, llama.cpp asserts every requested layer has a non-null res->t_layer_inp[il]. The kimi-k3 architecture graph builder (src/models/kimi-k3.cpp) never populates res->t_layer_inp[] β€” so the assert fires. Architectures that DO populate it (and thus support DSpark): deepseek4.cpp, bailingmoe3.cpp, gemma4.cpp.

Fix (applied on box /root/llama.cpp, backup at kimi-k3.cpp.bak): in the kimi-k3 graph builder layer loop, immediately after const auto & layer = model.layers[il];:

// expose the raw layer input for speculative draft (DSpark/dflash) feature taps
if ((size_t) il < cparams.embeddings_layer_inp.size() && cparams.embeddings_layer_inp[il]) {
    res->t_layer_inp[il] = inpL;
    cb(res->t_layer_inp[il], "layer_inp", il);
    ggml_build_forward_expand(gf, res->t_layer_inp[il]);
}

K3's inpL is already the plain per-layer input (no hyper-connection transform needed, unlike deepseek4's dsv4_hc_mean). Draft taps are all at il<93, so no post-loop tail extraction required. Rebuild: cmake --build . --config Release -j 56 --target llama-server. Result: DSpark decodes on K3 without crash. Upstream PR candidate.

1b. DSpark draft fails every step "invalid token[1] = -1" β†’ FIXED (GGUF mask token)

Symptom: After the t_layer_inp patch, the server loads + the main model decodes, but the draft fails EVERY step: init: invalid token[1] = -1 β†’ decode: failed to initialize batch β†’ llama_decode returned -1 β†’ draft: llama_decode returned -1. Main model still generates (falls back to no speculation) β†’ runs at baseline speed with extra overhead, zero spec speedup.

Root cause: DSpark builds the draft batch with a mask token for the multi-token block: common_batch_add(batch, i==0 ? dp.id_last : mask_token_id, ...) (speculative.cpp:1189). It resolves the mask id via llama_vocab_mask(vocab) which reads tokenizer.ggml.mask_token_id. Our rewritten draft GGUF copied K3's tokenizer keys (fix_draft_gguf.py) but K3 has NO mask token β†’ llama_vocab_mask() returns -1 β†’ batch fed token -1 β†’ embedding rejects it. The draft's TRAINED mask id lives in dflash.mask_token_id = 163824 (separate metadata key, present in both original and rewritten draft), but llama.cpp never reads it for the mask.

Fix: add tokenizer.ggml.mask_token_id = 163824 (uint32) to the draft GGUF KV section (add_mask_token.py, value taken from dflash.mask_token_id). Output draft_masked.gguf; repoint draft.gguf symlink at it. Log then shows mask_token_id=163824 and zero draft errors. (Arguably also a mainline improvement: fall back to dflash.mask_token_id when the vocab has no mask token.)

2. DSpark draft-model VRAM OOM β†’ fixed with -ngld 0

The 2.4GB Q8_0 DSpark draft tried to allocate a 4GB KV cache on GPU device 1, which is already near-full from the K3 trunk 8-way split. Fix: -ngld 0 (--gpu-layers-draft 0) runs the draft entirely on CPU. Draft is tiny; the main K3 trunk is the bottleneck anyway.

3. -fa on + large batch OOM on 16GB cards β†’ reduced batch + device-0 share

With --tensor-split 0.3,1,1,1,1,1,1,1 -b 512 -ub 512, enabling -fa on OOMs the compute pp buffers. Kitchen-sink config that reaches READY: --tensor-split 0.2,1,1,1,1,1,1,1 -b 256 -ub 256 -fa on. (Without -fa on, batch 512 + split 0.3 also OOMs compute pp buffers; the no-FA auto path was the previously-working config.)

4. Load-time OOM (RssFile 482GB→cgroup 503GB) → LLAMA_MMAP_NO_PREFETCH=1

src/llama-model.cpp:1663 calls ml.init_mappings(true, ...) β†’ MAP_POPULATE + MADV_WILLNEED eagerly faults all 1.5TB expert pages at load. Patch reads LLAMA_MMAP_NO_PREFETCH env β†’ init_mappings(!no_prefetch, ...). Result: RssFile 124GB at load, lazy LRU page-cache becomes the hot-expert cache (the architecture the user wants). This patch is REQUIRED on the home 768GB box too β€” stock llama.cpp cannot load Q4_K_XL.

5. HF download throttling β†’ presigned URL + aria2c

curl -L on huggingface.co/resolve is throttled to ~7KB/s on many datacenter routes. Resolve without -L, extract the presigned CloudFront URL, download with aria2c -x16 β†’ 255-281 MB/s. (lib/hf_direct.sh.)

6. Bash GROUPS is special β€” never use it as a var name

GROUPS is a bash builtin array (user's group IDs). Using it for suite group selection caused a silent no-op. Renamed to SEL.

7. The trunk is Q8_0, NOT 4-bit β€” and DSpark-on-CPU is a net loss (both fixed by Q4 trunk)

Discovery (user's instinct was right): "UD-Q4_K_XL" only 4-bit-quantizes the routed experts (MXFP4). The entire trunk β€” attention, shared experts (_shexp), output head, token embedding β€” is Q8_0 (8-bit), plus F32 norms. Measured across all 32 shards:

  • trunk Q8_0 = 59.6 GB (1116 tensors), norms F32 = 2.6 GB, experts = MXFP4 (rest of 1.4TB).

Why: K3 is QAT-trained in MXFP4 β€” the 4-bit experts ARE the reference model (no BF16 original). But the non-expert path stays higher-precision (activations MXFP8, non-expert weights higher precision). So Q8_0 trunk is a legit DOWN-quant from a higher-precision source, NOT an up-quant of 4-bit. Ref: dreaming.press "Kimi K3's Weights Are Already 4-Bit". Consequence: re-quantizing the TRUNK Q8_0β†’Q4_K is valid (source was >4-bit). NEVER re-quant the MXFP4 experts (destroys QAT calibration).

DSpark-on-CPU measured result: kitchen sink with -ngld 0 (draft on CPU) gave ks_cold tg 0.224 t/s vs ~0.5 t/s no-spec baseline = ~2x LOSS. Decode is pinned by CPU expert execution; the CPU draft forward steals the same 112 threads. DSpark only wins with the draft ON GPU, but the 58GB Q8_0 trunk fills all 8Γ—16GB cards. β†’ Need a smaller trunk.

The unlock (both home fit + GPU DSpark): requant trunk Q8_0β†’Q4_K.

  • Trunk 59.6β†’31.6 GB; GPU-resident 62.2β†’34.2 GB. Fits home 2Γ—3090 (48GB) AND frees ~28GB on the rental box β†’ DSpark draft can go on GPU.
  • llama-quantize CANNOT do this safely: --allow-requantize forces every non-overridden tensor (incl. MXFP4 experts) to the positional type β†’ dequant+requant experts β†’ QAT loss. Dry-run confirmed experts became q4_K/q6_K and total size GREW 1438771β†’1602481 MiB.
  • Solution: custom surgical rewriter requant_trunk.c (compiled on box at /root/k3-test/requant_trunk). Per-shard, split-in=split-out. Byte-copies MXFP4 experts
    • F32 norms unchanged; dequant Q8_0β†’F32 + requant F32β†’Q4_K (ggml dequantize_row_q8_0/quantize_row_q4_K_ref) for trunk tensors only. Falls back to copy for tensors whose dims[0] not divisible by 256 (e.g. attn_k_b/ssm_f_b at 128). Validated on shard 2: 43 requant + 62 copy, 47.5β†’44.9 GB, no crash.

Baselines (mainline, batch-1, 64in/64out unless noted)

  • Best known (t112, b512/ub512, split 0.3,1..., FA-auto/off, no spec): cold pp ~0.30-0.40 tg ~0.33-0.51; warm pp 2.5-2.8 tg 0.65-1.79 t/s.
  • DSpark-on-CPU kitchen sink: cold pp 0.237 tg 0.224 t/s (~2x WORSE than no-spec).
  • t56 regression: warm pp 0.62 tg ~0.49-0.59 β†’ t112 wins (SMT siblings help).
  • Decode is pinned by CPU expert execution, not cache warm-up (warm tg β‰ˆ cold tg).

Target architecture (home build)

  • 2Γ— 3090 24GB (48GB VRAM) + 768GB RAM + EPYC. Trunk Q8_0 58GB > 48GB β†’ Q4_K trunk requant (34.2GB incl. norms) fits with room for the DSpark draft on GPU. This is the chosen path (see #7). Untestable for exact 2Γ—3090 split on the 8Γ—16GB rental box, but the Q4_K trunk + GPU-draft DSpark combo IS testable there (frees ~28GB VRAM).
  • Experts: as many as fit in RAM (~150 hottest mlock-pinned, deferred), rest lazily faulted/evicted from SSD via kernel page-LRU (--cpu-moe + lazy mmap). Q4 only (no smaller quant β€” quality). Shared experts (_shexp) are GPU-resident, never evicted.
  • Spec (DSpark, now working) + expert offload combined = the realistic production case.