Instructions to use TessaCoil/K3-Stuff with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use TessaCoil/K3-Stuff with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf TessaCoil/K3-Stuff:Q8_0 # Run inference directly in the terminal: llama cli -hf TessaCoil/K3-Stuff:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf TessaCoil/K3-Stuff:Q8_0 # Run inference directly in the terminal: llama cli -hf TessaCoil/K3-Stuff:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf TessaCoil/K3-Stuff:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf TessaCoil/K3-Stuff:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf TessaCoil/K3-Stuff:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf TessaCoil/K3-Stuff:Q8_0
Use Docker
docker model run hf.co/TessaCoil/K3-Stuff:Q8_0
- LM Studio
- Jan
- Ollama
How to use TessaCoil/K3-Stuff with Ollama:
ollama run hf.co/TessaCoil/K3-Stuff:Q8_0
- Unsloth Desktop
- Pi
How to use TessaCoil/K3-Stuff with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TessaCoil/K3-Stuff:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TessaCoil/K3-Stuff:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use TessaCoil/K3-Stuff with Docker Model Runner:
docker model run hf.co/TessaCoil/K3-Stuff:Q8_0
- Lemonade
How to use TessaCoil/K3-Stuff with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull TessaCoil/K3-Stuff:Q8_0
Run and chat with the model
lemonade run user.K3-Stuff-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use TessaCoil/K3-Stuff with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TessaCoil/K3-Stuff:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TessaCoil/K3-Stuff:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TessaCoil/K3-Stuff with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TessaCoil/K3-Stuff:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TessaCoil/K3-Stuff:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
File size: 8,673 Bytes
ddf8c5b | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 122 123 124 125 126 127 128 129 130 131 132 133 134 135 136 137 138 139 140 141 142 143 | # K3 Rental Validation β Durable Findings & Fixes
Hard-won root causes, patches, and decisions. Transient state (PIDs, SHAs, run progress)
lives in the session log, not here.
## 1. K3 + DSpark speculative decoding crash β FIXED (patch)
**Symptom:** llama-server with `--spec-type draft-dspark -md draft.gguf` loads fine,
reaches READY, but crashes on the FIRST decode:
```
llama-graph.cpp:1376: GGML_ASSERT(t_layer_inp[il] != nullptr && "layer input tensor is null") failed
```
**Root cause:** DSpark/dflash speculative decoding extracts intermediate "layer input"
features from the target (K3) model to feed the draft. It taps specific layers via
`cparams.embeddings_layer_inp[]` (the Lucebox draft GGUF requests `target_layer_ids =
[7, 23, 51, 67, 83]`). At `set_outputs()` time, llama.cpp asserts every requested layer
has a non-null `res->t_layer_inp[il]`. The kimi-k3 architecture graph builder
(`src/models/kimi-k3.cpp`) **never populates `res->t_layer_inp[]`** β so the assert fires.
Architectures that DO populate it (and thus support DSpark): `deepseek4.cpp`,
`bailingmoe3.cpp`, `gemma4.cpp`.
**Fix (applied on box `/root/llama.cpp`, backup at `kimi-k3.cpp.bak`):** in the kimi-k3
graph builder layer loop, immediately after `const auto & layer = model.layers[il];`:
```cpp
// expose the raw layer input for speculative draft (DSpark/dflash) feature taps
if ((size_t) il < cparams.embeddings_layer_inp.size() && cparams.embeddings_layer_inp[il]) {
res->t_layer_inp[il] = inpL;
cb(res->t_layer_inp[il], "layer_inp", il);
ggml_build_forward_expand(gf, res->t_layer_inp[il]);
}
```
K3's `inpL` is already the plain per-layer input (no hyper-connection transform needed,
unlike deepseek4's `dsv4_hc_mean`). Draft taps are all at il<93, so no post-loop tail
extraction required. Rebuild: `cmake --build . --config Release -j 56 --target llama-server`.
Result: DSpark decodes on K3 without crash. **Upstream PR candidate.**
## 1b. DSpark draft fails every step "invalid token[1] = -1" β FIXED (GGUF mask token)
**Symptom:** After the t_layer_inp patch, the server loads + the main model decodes, but
the draft fails EVERY step: `init: invalid token[1] = -1` β `decode: failed to initialize
batch` β `llama_decode returned -1` β `draft: llama_decode returned -1`. Main model still
generates (falls back to no speculation) β runs at baseline speed with extra overhead,
zero spec speedup.
**Root cause:** DSpark builds the draft batch with a mask token for the multi-token block:
`common_batch_add(batch, i==0 ? dp.id_last : mask_token_id, ...)` (speculative.cpp:1189).
It resolves the mask id via `llama_vocab_mask(vocab)` which reads
`tokenizer.ggml.mask_token_id`. Our rewritten draft GGUF copied K3's tokenizer keys
(`fix_draft_gguf.py`) but K3 has NO mask token β `llama_vocab_mask()` returns -1 β batch
fed token -1 β embedding rejects it. The draft's TRAINED mask id lives in
`dflash.mask_token_id` = **163824** (separate metadata key, present in both original and
rewritten draft), but llama.cpp never reads it for the mask.
**Fix:** add `tokenizer.ggml.mask_token_id = 163824` (uint32) to the draft GGUF KV section
(`add_mask_token.py`, value taken from `dflash.mask_token_id`). Output
`draft_masked.gguf`; repoint `draft.gguf` symlink at it. Log then shows
`mask_token_id=163824` and zero draft errors. (Arguably also a mainline improvement:
fall back to `dflash.mask_token_id` when the vocab has no mask token.)
## 2. DSpark draft-model VRAM OOM β fixed with `-ngld 0`
The 2.4GB Q8_0 DSpark draft tried to allocate a 4GB KV cache on GPU device 1, which is
already near-full from the K3 trunk 8-way split. Fix: `-ngld 0` (`--gpu-layers-draft 0`)
runs the draft entirely on CPU. Draft is tiny; the main K3 trunk is the bottleneck anyway.
## 3. `-fa on` + large batch OOM on 16GB cards β reduced batch + device-0 share
With `--tensor-split 0.3,1,1,1,1,1,1,1 -b 512 -ub 512`, enabling `-fa on` OOMs the
compute pp buffers. Kitchen-sink config that reaches READY: `--tensor-split
0.2,1,1,1,1,1,1,1 -b 256 -ub 256 -fa on`. (Without `-fa on`, batch 512 + split 0.3 also
OOMs compute pp buffers; the no-FA auto path was the previously-working config.)
## 4. Load-time OOM (RssFile 482GBβcgroup 503GB) β `LLAMA_MMAP_NO_PREFETCH=1`
`src/llama-model.cpp:1663` calls `ml.init_mappings(true, ...)` β MAP_POPULATE +
MADV_WILLNEED eagerly faults all 1.5TB expert pages at load. Patch reads
`LLAMA_MMAP_NO_PREFETCH` env β `init_mappings(!no_prefetch, ...)`. Result: RssFile 124GB
at load, lazy LRU page-cache becomes the hot-expert cache (the architecture the user wants).
This patch is REQUIRED on the home 768GB box too β stock llama.cpp cannot load Q4_K_XL.
## 5. HF download throttling β presigned URL + aria2c
`curl -L` on huggingface.co/resolve is throttled to ~7KB/s on many datacenter routes.
Resolve without `-L`, extract the presigned CloudFront URL, download with `aria2c -x16`
β 255-281 MB/s. (`lib/hf_direct.sh`.)
## 6. Bash `GROUPS` is special β never use it as a var name
`GROUPS` is a bash builtin array (user's group IDs). Using it for suite group selection
caused a silent no-op. Renamed to `SEL`.
## 7. The trunk is Q8_0, NOT 4-bit β and DSpark-on-CPU is a net loss (both fixed by Q4 trunk)
**Discovery (user's instinct was right):** "UD-Q4_K_XL" only 4-bit-quantizes the **routed
experts** (MXFP4). The entire **trunk β attention, shared experts (`_shexp`), output head,
token embedding β is Q8_0** (8-bit), plus F32 norms. Measured across all 32 shards:
- trunk Q8_0 = 59.6 GB (1116 tensors), norms F32 = 2.6 GB, experts = MXFP4 (rest of 1.4TB).
**Why:** K3 is QAT-trained in MXFP4 β the 4-bit experts ARE the reference model (no BF16
original). But the **non-expert path stays higher-precision** (activations MXFP8, non-expert
weights higher precision). So Q8_0 trunk is a legit DOWN-quant from a higher-precision
source, NOT an up-quant of 4-bit. Ref: dreaming.press "Kimi K3's Weights Are Already 4-Bit".
Consequence: re-quantizing the TRUNK Q8_0βQ4_K is valid (source was >4-bit). NEVER re-quant
the MXFP4 experts (destroys QAT calibration).
**DSpark-on-CPU measured result:** kitchen sink with `-ngld 0` (draft on CPU) gave
**ks_cold tg 0.224 t/s vs ~0.5 t/s no-spec baseline = ~2x LOSS**. Decode is pinned by CPU
expert execution; the CPU draft forward steals the same 112 threads. DSpark only wins with
the draft ON GPU, but the 58GB Q8_0 trunk fills all 8Γ16GB cards. β Need a smaller trunk.
**The unlock (both home fit + GPU DSpark):** requant trunk Q8_0βQ4_K.
- Trunk 59.6β31.6 GB; GPU-resident 62.2β34.2 GB. Fits home 2Γ3090 (48GB) AND frees ~28GB
on the rental box β DSpark draft can go on GPU.
- **llama-quantize CANNOT do this safely**: `--allow-requantize` forces every non-overridden
tensor (incl. MXFP4 experts) to the positional type β dequant+requant experts β QAT loss.
Dry-run confirmed experts became q4_K/q6_K and total size GREW 1438771β1602481 MiB.
- **Solution: custom surgical rewriter `requant_trunk.c`** (compiled on box at
`/root/k3-test/requant_trunk`). Per-shard, split-in=split-out. Byte-copies MXFP4 experts
+ F32 norms unchanged; dequant Q8_0βF32 + requant F32βQ4_K (ggml
`dequantize_row_q8_0`/`quantize_row_q4_K_ref`) for trunk tensors only. Falls back to copy
for tensors whose dims[0] not divisible by 256 (e.g. attn_k_b/ssm_f_b at 128).
Validated on shard 2: 43 requant + 62 copy, 47.5β44.9 GB, no crash.
## Baselines (mainline, batch-1, 64in/64out unless noted)
- Best known (t112, b512/ub512, split 0.3,1..., FA-auto/off, no spec): cold pp ~0.30-0.40
tg ~0.33-0.51; warm pp 2.5-2.8 tg 0.65-1.79 t/s.
- DSpark-on-CPU kitchen sink: cold pp 0.237 tg 0.224 t/s (~2x WORSE than no-spec).
- t56 regression: warm pp 0.62 tg ~0.49-0.59 β **t112 wins** (SMT siblings help).
- Decode is pinned by CPU expert execution, not cache warm-up (warm tg β cold tg).
## Target architecture (home build)
- 2Γ 3090 24GB (48GB VRAM) + 768GB RAM + EPYC. Trunk Q8_0 58GB > 48GB β **Q4_K trunk
requant (34.2GB incl. norms) fits with room for the DSpark draft on GPU.** This is the
chosen path (see #7). Untestable for exact 2Γ3090 split on the 8Γ16GB rental box, but
the Q4_K trunk + GPU-draft DSpark combo IS testable there (frees ~28GB VRAM).
- Experts: as many as fit in RAM (~150 hottest mlock-pinned, deferred), rest lazily
faulted/evicted from SSD via kernel page-LRU (`--cpu-moe` + lazy mmap). Q4 only
(no smaller quant β quality). Shared experts (`_shexp`) are GPU-resident, never evicted.
- Spec (DSpark, now working) + expert offload combined = the realistic production case.
|