Text Generation
GGUF
llama.cpp
agent
coding
reasoning
tool-use
function-calling
quantized
cuda
metal
conversational
Instructions to use badtheorylabs/BTL-3-Compact with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use badtheorylabs/BTL-3-Compact with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./llama-cli -hf badtheorylabs/BTL-3-Compact
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./build/bin/llama-cli -hf badtheorylabs/BTL-3-Compact
Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- LM Studio
- Jan
- vLLM
How to use badtheorylabs/BTL-3-Compact with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "badtheorylabs/BTL-3-Compact" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badtheorylabs/BTL-3-Compact", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Ollama
How to use badtheorylabs/BTL-3-Compact with Ollama:
ollama run hf.co/badtheorylabs/BTL-3-Compact
- Unsloth Studio
How to use badtheorylabs/BTL-3-Compact with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
- Pi
How to use badtheorylabs/BTL-3-Compact with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "badtheorylabs/BTL-3-Compact" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use badtheorylabs/BTL-3-Compact with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default badtheorylabs/BTL-3-Compact
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use badtheorylabs/BTL-3-Compact with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "badtheorylabs/BTL-3-Compact" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use badtheorylabs/BTL-3-Compact with Docker Model Runner:
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Lemonade
How to use badtheorylabs/BTL-3-Compact with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull badtheorylabs/BTL-3-Compact
Run and chat with the model
lemonade run user.BTL-3-Compact-{{QUANT_TAG}}List all available models
lemonade list
File size: 2,459 Bytes
274ffba 1ccb6ee 274ffba 1ccb6ee 274ffba 1ccb6ee 274ffba 1ccb6ee 274ffba | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 | {
"schema_version": 1,
"release": "BTL-3 Compact — Thinking Escape r8 step160 v3 canonical refresh",
"checkpoint": "RL-0013",
"architecture": "Qwen3.6-27B",
"model": {
"path": "model/BTL-3-Compact-AVQ2.gguf",
"bytes": 8392369600,
"sha256": "0a4d9ddee49e5aa93586a792bd4d452ea837229d49d22e54212dde87a5c9888a"
},
"runtimes": {
"supported": [
{
"name": "BTL-3-Compact-macos-arm64",
"path": "runtimes/supported/BTL-3-Compact-macos-arm64",
"bundle": "BTL-3 Compact macOS arm64",
"status": "verified native runtime"
}
],
"preview": [
{
"name": "BTL-3-Compact-linux-arm64-cuda",
"path": "runtimes/preview/BTL-3-Compact-linux-arm64-cuda",
"bundle": "BTL-3 Compact Linux arm64 CUDA (DGX Spark)",
"status": "cross-compiled; NVIDIA runtime conformance pending"
}
]
},
"integrations": [
"integrations/btl3-native",
"integrations/ollama"
],
"documentation": [
"docs/launch-btl3-compact.md",
"docs/launch-btl3-cuda.md",
"docs/patched-ollama-and-lmstudio.md",
"docs/btl-3-compact-gguf-exporter.md"
],
"evidence": [
"evidence/BTL-3-Compact-AVQ2.report.json",
"evidence/native-load-report.json",
"evidence/rtx-pro-6000-speed-2026-07-19.md",
"evidence/compact-validation.md",
"evidence/thinking-escape-v3-smoke.json",
"evidence/thinking-escape-v3-tool-gate.json",
"evidence/thinking-escape-v3-native-runtime.json",
"evidence/thinking-escape-v3-training-manifest.json"
],
"tools": [
"tools/install_consumer_bundle.py"
],
"licenses": [
"licenses/LICENSE",
"licenses/LICENSE.runtime",
"licenses/THIRD_PARTY_NOTICES.md"
],
"stock_ollama_compatible": false,
"stock_lm_studio_engine_compatible": false,
"default_thinking": false,
"canonical_refresh": {
"date": "2026-07-24",
"behavior_adapter": {
"identity": "thinking-escape-r8-step160-v3",
"rank": 8,
"alpha": 16,
"tensor_count": 132,
"payload_bytes": 32440320,
"sha256": "7ba9f80be97287775575022e84e735ac3c38b343c0bb83c5c3b0a86c53d6c881"
},
"non_behavior_tensors_verified_unchanged": 2284,
"packed_decoder_unchanged": true,
"rank32_output_head_correction_unchanged": true,
"tool_retention": "62/63",
"thinking_final_answer_rate": "5/12",
"thinking_coding_termination": "0/3",
"approved_as_full_thinking_fix": false
}
}
|