How to use from
llama.cpp
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Snapkitty/snapkitty-harness:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Snapkitty/snapkitty-harness:Q4_K_M
Install from WinGet (Windows)
winget install llama.cpp
# Start a local OpenAI-compatible server with a web UI:
llama serve -hf Snapkitty/snapkitty-harness:Q4_K_M
# Run inference directly in the terminal:
llama cli -hf Snapkitty/snapkitty-harness:Q4_K_M
Use pre-built binary
# Download pre-built binary from:
# https://github.com/ggerganov/llama.cpp/releases
# Start a local OpenAI-compatible server with a web UI:
./llama-server -hf Snapkitty/snapkitty-harness:Q4_K_M
# Run inference directly in the terminal:
./llama-cli -hf Snapkitty/snapkitty-harness:Q4_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git
cd llama.cpp
cmake -B build
cmake --build build -j --target llama-server llama-cli
# Start a local OpenAI-compatible server with a web UI:
./build/bin/llama-server -hf Snapkitty/snapkitty-harness:Q4_K_M
# Run inference directly in the terminal:
./build/bin/llama-cli -hf Snapkitty/snapkitty-harness:Q4_K_M
Use Docker
docker model run hf.co/Snapkitty/snapkitty-harness:Q4_K_M
Quick Links

snapkitty-harness

Reference model for the SnapKitty harness: NVIDIA Nemotron-Mini-4B-Instruct as a 2.7 GB Q4_K_M GGUF, ready for Ollama and llama.cpp. Pair it with snapkitty-nemotron-harness and benchmark tool gating with ToolGate-Bench.

Part of the SNAPKITTYWEST Sovereign Compute constellation.

Unified theory: 10.5281/zenodo.21816366

Research Papers → GitHub →

Model details

Property Value
Base model nvidia/Nemotron-Mini-4B-Instruct (quantized)
Architecture Nemotron (4B parameters)
Format GGUF v3, Q4_K_M (4-bit)
Context length 4,096 tokens
Layers / hidden size 32 / 3,072
Attention 24 heads, 8 KV heads (GQA)
Feed-forward size 9,216
Vocabulary 256,000 tokens
RoPE base 10000, 64 dims
Chat template Nemotron (<extra_id_0>System, <extra_id_1>User, <extra_id_1>Assistant)
File snapkitty-harness.Q4_K_M.gguf, 2.70 GB
SHA-256 1dcbd925825b41744ddc2fc3047db6d3ad0aecf8d336f4fadc044eaaf79779d5

Run it locally

Ollama

ollama run hf.co/Snapkitty/snapkitty-harness:Q4_K_M

llama.cpp

llama-cli -hf Snapkitty/snapkitty-harness:Q4_K_M -cnv
# or a local file
llama-server -m snapkitty-harness.Q4_K_M.gguf -c 4096

LM Studio: search for Snapkitty/snapkitty-harness and pick the Q4_K_M file.

Python (llama-cpp-python)

from llama_cpp import Llama
llm = Llama.from_pretrained(repo_id="Snapkitty/snapkitty-harness", filename="snapkitty-harness.Q4_K_M.gguf", n_ctx=4096)
out = llm.create_chat_completion(messages=[{"role": "user", "content": "Explain SUBLEQ in one paragraph."}])
print(out["choices"][0]["message"]["content"])

Verify the download

sha256sum snapkitty-harness.Q4_K_M.gguf
# 1dcbd925825b41744ddc2fc3047db6d3ad0aecf8d336f4fadc044eaaf79779d5

Citation

@misc{snapkitty_snapkitty_harness_2026,
  title        = {snapkitty-harness: Nemotron 4B GGUF (Q4\_K\_M)},
  author       = {{Snapkitty Collective LLC}},
  year         = {2026},
  howpublished = {\url{https://huggingface.co/Snapkitty/snapkitty-harness}}
}

Please also cite the base model, NVIDIA Nemotron-Mini-4B-Instruct.

License

The model weights are a derivative of NVIDIA Nemotron-Mini-4B-Instruct and are distributed under the NVIDIA Community Model License. See LICENSE.

Downloads last month
40
GGUF
Model size
4B params
Architecture
nemotron
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Snapkitty/snapkitty-harness

Quantized
(29)
this model

Space using Snapkitty/snapkitty-harness 1

Collection including Snapkitty/snapkitty-harness