Instructions to use kambrosius/veritas-coder-7b-gguf with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use kambrosius/veritas-coder-7b-gguf with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M # Run inference directly in the terminal: llama cli -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M # Run inference directly in the terminal: llama cli -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M # Run inference directly in the terminal: ./llama-cli -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M # Run inference directly in the terminal: ./build/bin/llama-cli -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M
Use Docker
docker model run hf.co/kambrosius/veritas-coder-7b-gguf:Q5_K_M
- LM Studio
- Jan
- vLLM
How to use kambrosius/veritas-coder-7b-gguf with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kambrosius/veritas-coder-7b-gguf" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kambrosius/veritas-coder-7b-gguf", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/kambrosius/veritas-coder-7b-gguf:Q5_K_M
- Ollama
How to use kambrosius/veritas-coder-7b-gguf with Ollama:
ollama run hf.co/kambrosius/veritas-coder-7b-gguf:Q5_K_M
- Unsloth Studio
How to use kambrosius/veritas-coder-7b-gguf with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kambrosius/veritas-coder-7b-gguf to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for kambrosius/veritas-coder-7b-gguf to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for kambrosius/veritas-coder-7b-gguf to start chatting
- Pi
How to use kambrosius/veritas-coder-7b-gguf with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "kambrosius/veritas-coder-7b-gguf:Q5_K_M" } ] } } }Run Pi
# Start Pi in your project directory: pi
- OpenClaw new
How to use kambrosius/veritas-coder-7b-gguf with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "kambrosius/veritas-coder-7b-gguf:Q5_K_M" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use kambrosius/veritas-coder-7b-gguf with Docker Model Runner:
docker model run hf.co/kambrosius/veritas-coder-7b-gguf:Q5_K_M
- Lemonade
How to use kambrosius/veritas-coder-7b-gguf with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull kambrosius/veritas-coder-7b-gguf:Q5_K_M
Run and chat with the model
lemonade run user.veritas-coder-7b-gguf-Q5_K_M
List all available models
lemonade list
- Hermes Agent
How to use kambrosius/veritas-coder-7b-gguf with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default kambrosius/veritas-coder-7b-gguf:Q5_K_M
Run Hermes
hermes
- Atomic Chat
veritas-coder-7b (GGUF)
The language "mind" of the veritas-hive unified system: Qwen2.5-Coder-7B-Instruct, QLoRA-distilled
with the veritas epistemic discipline (recall-before-generate, seal only tool-origin evidence,
promote a claim only on sealed evidence), plus 25 epistemic control tokens.
Honest scope. This repo is the mind only. veritas-hive's soundness guarantee lives in a deterministic ledger runtime (hash-chain + HMAC seals + trust calculus + the MCP tools) that ships alongside the weights — not inside them. The model can be wrong; the ledger is what stops a guess from becoming a belief. GGUF cannot represent the ledger or the GWM world-model (
.pt); those run as separate processes. See the veritas-hive repo for the full offline system.
Files
| file | quant | size | sha256 |
|---|---|---|---|
veritas-coder-7b.Q5_K_M.gguf |
Q5_K_M | 5.44 GB | 2756619586bc2fd2a585de5ade9cee86828a4ad8112469e9b7a3ca8438c2a335 |
Run with llama.cpp
llama-cli -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M -p "Write a Python LRU cache."
llama-server -hf kambrosius/veritas-coder-7b-gguf:Q5_K_M --port 8080
Run with Ollama
ollama run hf.co/kambrosius/veritas-coder-7b-gguf:Q5_K_M
This repo is private. For Ollama/llama.cpp to pull it, either make the repo public, or add your Hugging Face access token / SSH key to the client (Ollama: add your HF SSH public key at https://huggingface.co/settings/keys).
Prompt format
ChatML (Qwen2.5). A veritas system prompt is recommended:
You are veritas-hive. Recall before you generate. Prefer verified facts over guesses.
Never assert a sealed/promoted claim you cannot back with evidence — hedge instead.
State uncertainty plainly.
Full build pipeline
This repo ships the complete, reproducible process — see PIPELINE.md:
training/ (data → QLoRA SFT → merge → GGUF export → eval, with real results in training/out/),
llama-veritas/ (the custom veritas-arch llama.cpp fork), runtime/ (Ollama Modelfile + the
model⇄MCP bridge), and docs/ (architecture, whitepaper, threat model).
Provenance
- Base: Qwen/Qwen2.5-Coder-7B-Instruct
- Adaptation: QLoRA SFT — verified code/math correctness + self-correction traces + epistemic control-token protocol; trained on an RTX 5090 (sm_120, torch cu128).
- Quantization: llama.cpp
llama-quantize→ Q5_K_M. - Eval: correctness measured through real test execution (dojo / EvalPlus / BigCodeBench harnesses),
not self-report. See the veritas-hive repo
training/out/for numbers.
Limitations
- 7B: broad capability comes from the base model; this distill imprints discipline, not raw IQ.
- Tool-calling: as a code-tuned model it sometimes emits tool calls as code/JSON rather than the structured format — the veritas-local bridge parses both.
- Not a substitute for the ledger: run the MCP runtime alongside for the actual epistemic guarantees.
- Downloads last month
- -
5-bit