Instructions to use TessaCoil/K3-Stuff with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use TessaCoil/K3-Stuff with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf TessaCoil/K3-Stuff:Q8_0 # Run inference directly in the terminal: llama cli -hf TessaCoil/K3-Stuff:Q8_0
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf TessaCoil/K3-Stuff:Q8_0 # Run inference directly in the terminal: llama cli -hf TessaCoil/K3-Stuff:Q8_0
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf TessaCoil/K3-Stuff:Q8_0 # Run inference directly in the terminal: ./llama-cli -hf TessaCoil/K3-Stuff:Q8_0
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf TessaCoil/K3-Stuff:Q8_0 # Run inference directly in the terminal: ./build/bin/llama-cli -hf TessaCoil/K3-Stuff:Q8_0
Use Docker
docker model run hf.co/TessaCoil/K3-Stuff:Q8_0
- LM Studio
- Jan
- Ollama
How to use TessaCoil/K3-Stuff with Ollama:
ollama run hf.co/TessaCoil/K3-Stuff:Q8_0
- Unsloth Desktop
- Pi
How to use TessaCoil/K3-Stuff with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TessaCoil/K3-Stuff:Q8_0
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "TessaCoil/K3-Stuff:Q8_0" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use TessaCoil/K3-Stuff with Docker Model Runner:
docker model run hf.co/TessaCoil/K3-Stuff:Q8_0
- Lemonade
How to use TessaCoil/K3-Stuff with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull TessaCoil/K3-Stuff:Q8_0
Run and chat with the model
lemonade run user.K3-Stuff-Q8_0
List all available models
lemonade list
- Hermes Agent
How to use TessaCoil/K3-Stuff with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TessaCoil/K3-Stuff:Q8_0
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default TessaCoil/K3-Stuff:Q8_0
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use TessaCoil/K3-Stuff with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf TessaCoil/K3-Stuff:Q8_0
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "TessaCoil/K3-Stuff:Q8_0" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
K3 Rental Validation β Durable Findings & Fixes
Hard-won root causes, patches, and decisions. Transient state (PIDs, SHAs, run progress) lives in the session log, not here.
1. K3 + DSpark speculative decoding crash β FIXED (patch)
Symptom: llama-server with --spec-type draft-dspark -md draft.gguf loads fine,
reaches READY, but crashes on the FIRST decode:
llama-graph.cpp:1376: GGML_ASSERT(t_layer_inp[il] != nullptr && "layer input tensor is null") failed
Root cause: DSpark/dflash speculative decoding extracts intermediate "layer input"
features from the target (K3) model to feed the draft. It taps specific layers via
cparams.embeddings_layer_inp[] (the Lucebox draft GGUF requests target_layer_ids = [7, 23, 51, 67, 83]). At set_outputs() time, llama.cpp asserts every requested layer
has a non-null res->t_layer_inp[il]. The kimi-k3 architecture graph builder
(src/models/kimi-k3.cpp) never populates res->t_layer_inp[] β so the assert fires.
Architectures that DO populate it (and thus support DSpark): deepseek4.cpp,
bailingmoe3.cpp, gemma4.cpp.
Fix (applied on box /root/llama.cpp, backup at kimi-k3.cpp.bak): in the kimi-k3
graph builder layer loop, immediately after const auto & layer = model.layers[il];:
// expose the raw layer input for speculative draft (DSpark/dflash) feature taps
if ((size_t) il < cparams.embeddings_layer_inp.size() && cparams.embeddings_layer_inp[il]) {
res->t_layer_inp[il] = inpL;
cb(res->t_layer_inp[il], "layer_inp", il);
ggml_build_forward_expand(gf, res->t_layer_inp[il]);
}
K3's inpL is already the plain per-layer input (no hyper-connection transform needed,
unlike deepseek4's dsv4_hc_mean). Draft taps are all at il<93, so no post-loop tail
extraction required. Rebuild: cmake --build . --config Release -j 56 --target llama-server.
Result: DSpark decodes on K3 without crash. Upstream PR candidate.
1b. DSpark draft fails every step "invalid token[1] = -1" β FIXED (GGUF mask token)
Symptom: After the t_layer_inp patch, the server loads + the main model decodes, but
the draft fails EVERY step: init: invalid token[1] = -1 β decode: failed to initialize batch β llama_decode returned -1 β draft: llama_decode returned -1. Main model still
generates (falls back to no speculation) β runs at baseline speed with extra overhead,
zero spec speedup.
Root cause: DSpark builds the draft batch with a mask token for the multi-token block:
common_batch_add(batch, i==0 ? dp.id_last : mask_token_id, ...) (speculative.cpp:1189).
It resolves the mask id via llama_vocab_mask(vocab) which reads
tokenizer.ggml.mask_token_id. Our rewritten draft GGUF copied K3's tokenizer keys
(fix_draft_gguf.py) but K3 has NO mask token β llama_vocab_mask() returns -1 β batch
fed token -1 β embedding rejects it. The draft's TRAINED mask id lives in
dflash.mask_token_id = 163824 (separate metadata key, present in both original and
rewritten draft), but llama.cpp never reads it for the mask.
Fix: add tokenizer.ggml.mask_token_id = 163824 (uint32) to the draft GGUF KV section
(add_mask_token.py, value taken from dflash.mask_token_id). Output
draft_masked.gguf; repoint draft.gguf symlink at it. Log then shows
mask_token_id=163824 and zero draft errors. (Arguably also a mainline improvement:
fall back to dflash.mask_token_id when the vocab has no mask token.)
2. DSpark draft-model VRAM OOM β fixed with -ngld 0
The 2.4GB Q8_0 DSpark draft tried to allocate a 4GB KV cache on GPU device 1, which is
already near-full from the K3 trunk 8-way split. Fix: -ngld 0 (--gpu-layers-draft 0)
runs the draft entirely on CPU. Draft is tiny; the main K3 trunk is the bottleneck anyway.
3. -fa on + large batch OOM on 16GB cards β reduced batch + device-0 share
With --tensor-split 0.3,1,1,1,1,1,1,1 -b 512 -ub 512, enabling -fa on OOMs the
compute pp buffers. Kitchen-sink config that reaches READY: --tensor-split 0.2,1,1,1,1,1,1,1 -b 256 -ub 256 -fa on. (Without -fa on, batch 512 + split 0.3 also
OOMs compute pp buffers; the no-FA auto path was the previously-working config.)
4. Load-time OOM (RssFile 482GBβcgroup 503GB) β LLAMA_MMAP_NO_PREFETCH=1
src/llama-model.cpp:1663 calls ml.init_mappings(true, ...) β MAP_POPULATE +
MADV_WILLNEED eagerly faults all 1.5TB expert pages at load. Patch reads
LLAMA_MMAP_NO_PREFETCH env β init_mappings(!no_prefetch, ...). Result: RssFile 124GB
at load, lazy LRU page-cache becomes the hot-expert cache (the architecture the user wants).
This patch is REQUIRED on the home 768GB box too β stock llama.cpp cannot load Q4_K_XL.
5. HF download throttling β presigned URL + aria2c
curl -L on huggingface.co/resolve is throttled to ~7KB/s on many datacenter routes.
Resolve without -L, extract the presigned CloudFront URL, download with aria2c -x16
β 255-281 MB/s. (lib/hf_direct.sh.)
6. Bash GROUPS is special β never use it as a var name
GROUPS is a bash builtin array (user's group IDs). Using it for suite group selection
caused a silent no-op. Renamed to SEL.
7. The trunk is Q8_0, NOT 4-bit β and DSpark-on-CPU is a net loss (both fixed by Q4 trunk)
Discovery (user's instinct was right): "UD-Q4_K_XL" only 4-bit-quantizes the routed
experts (MXFP4). The entire trunk β attention, shared experts (_shexp), output head,
token embedding β is Q8_0 (8-bit), plus F32 norms. Measured across all 32 shards:
- trunk Q8_0 = 59.6 GB (1116 tensors), norms F32 = 2.6 GB, experts = MXFP4 (rest of 1.4TB).
Why: K3 is QAT-trained in MXFP4 β the 4-bit experts ARE the reference model (no BF16 original). But the non-expert path stays higher-precision (activations MXFP8, non-expert weights higher precision). So Q8_0 trunk is a legit DOWN-quant from a higher-precision source, NOT an up-quant of 4-bit. Ref: dreaming.press "Kimi K3's Weights Are Already 4-Bit". Consequence: re-quantizing the TRUNK Q8_0βQ4_K is valid (source was >4-bit). NEVER re-quant the MXFP4 experts (destroys QAT calibration).
DSpark-on-CPU measured result: kitchen sink with -ngld 0 (draft on CPU) gave
ks_cold tg 0.224 t/s vs ~0.5 t/s no-spec baseline = ~2x LOSS. Decode is pinned by CPU
expert execution; the CPU draft forward steals the same 112 threads. DSpark only wins with
the draft ON GPU, but the 58GB Q8_0 trunk fills all 8Γ16GB cards. β Need a smaller trunk.
The unlock (both home fit + GPU DSpark): requant trunk Q8_0βQ4_K.
- Trunk 59.6β31.6 GB; GPU-resident 62.2β34.2 GB. Fits home 2Γ3090 (48GB) AND frees ~28GB on the rental box β DSpark draft can go on GPU.
- llama-quantize CANNOT do this safely:
--allow-requantizeforces every non-overridden tensor (incl. MXFP4 experts) to the positional type β dequant+requant experts β QAT loss. Dry-run confirmed experts became q4_K/q6_K and total size GREW 1438771β1602481 MiB. - Solution: custom surgical rewriter
requant_trunk.c(compiled on box at/root/k3-test/requant_trunk). Per-shard, split-in=split-out. Byte-copies MXFP4 experts- F32 norms unchanged; dequant Q8_0βF32 + requant F32βQ4_K (ggml
dequantize_row_q8_0/quantize_row_q4_K_ref) for trunk tensors only. Falls back to copy for tensors whose dims[0] not divisible by 256 (e.g. attn_k_b/ssm_f_b at 128). Validated on shard 2: 43 requant + 62 copy, 47.5β44.9 GB, no crash.
- F32 norms unchanged; dequant Q8_0βF32 + requant F32βQ4_K (ggml
Baselines (mainline, batch-1, 64in/64out unless noted)
- Best known (t112, b512/ub512, split 0.3,1..., FA-auto/off, no spec): cold pp ~0.30-0.40 tg ~0.33-0.51; warm pp 2.5-2.8 tg 0.65-1.79 t/s.
- DSpark-on-CPU kitchen sink: cold pp 0.237 tg 0.224 t/s (~2x WORSE than no-spec).
- t56 regression: warm pp 0.62 tg ~0.49-0.59 β t112 wins (SMT siblings help).
- Decode is pinned by CPU expert execution, not cache warm-up (warm tg β cold tg).
Target architecture (home build)
- 2Γ 3090 24GB (48GB VRAM) + 768GB RAM + EPYC. Trunk Q8_0 58GB > 48GB β Q4_K trunk requant (34.2GB incl. norms) fits with room for the DSpark draft on GPU. This is the chosen path (see #7). Untestable for exact 2Γ3090 split on the 8Γ16GB rental box, but the Q4_K trunk + GPU-draft DSpark combo IS testable there (frees ~28GB VRAM).
- Experts: as many as fit in RAM (~150 hottest mlock-pinned, deferred), rest lazily
faulted/evicted from SSD via kernel page-LRU (
--cpu-moe+ lazy mmap). Q4 only (no smaller quant β quality). Shared experts (_shexp) are GPU-resident, never evicted. - Spec (DSpark, now working) + expert offload combined = the realistic production case.