Text Generation
GGUF
llama.cpp
agent
coding
reasoning
tool-use
function-calling
quantized
cuda
metal
conversational
Instructions to use badtheorylabs/BTL-3-Compact with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use badtheorylabs/BTL-3-Compact with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./llama-cli -hf badtheorylabs/BTL-3-Compact
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./build/bin/llama-cli -hf badtheorylabs/BTL-3-Compact
Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- LM Studio
- Jan
- vLLM
How to use badtheorylabs/BTL-3-Compact with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "badtheorylabs/BTL-3-Compact" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badtheorylabs/BTL-3-Compact", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Ollama
How to use badtheorylabs/BTL-3-Compact with Ollama:
ollama run hf.co/badtheorylabs/BTL-3-Compact
- Unsloth Studio
How to use badtheorylabs/BTL-3-Compact with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
- Pi
How to use badtheorylabs/BTL-3-Compact with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "badtheorylabs/BTL-3-Compact" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use badtheorylabs/BTL-3-Compact with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default badtheorylabs/BTL-3-Compact
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use badtheorylabs/BTL-3-Compact with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "badtheorylabs/BTL-3-Compact" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use badtheorylabs/BTL-3-Compact with Docker Model Runner:
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Lemonade
How to use badtheorylabs/BTL-3-Compact with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull badtheorylabs/BTL-3-Compact
Run and chat with the model
lemonade run user.BTL-3-Compact-{{QUANT_TAG}}List all available models
lemonade list
| { | |
| "schema_version": 1, | |
| "release": "BTL-3 Compact — Thinking Escape r8 step160 v3 canonical refresh", | |
| "checkpoint": "RL-0013", | |
| "architecture": "Qwen3.6-27B", | |
| "model": { | |
| "path": "model/BTL-3-Compact-AVQ2.gguf", | |
| "bytes": 8392369600, | |
| "sha256": "0a4d9ddee49e5aa93586a792bd4d452ea837229d49d22e54212dde87a5c9888a" | |
| }, | |
| "runtimes": { | |
| "supported": [ | |
| { | |
| "name": "BTL-3-Compact-macos-arm64", | |
| "path": "runtimes/supported/BTL-3-Compact-macos-arm64", | |
| "bundle": "BTL-3 Compact macOS arm64", | |
| "status": "verified native runtime" | |
| } | |
| ], | |
| "preview": [ | |
| { | |
| "name": "BTL-3-Compact-linux-arm64-cuda", | |
| "path": "runtimes/preview/BTL-3-Compact-linux-arm64-cuda", | |
| "bundle": "BTL-3 Compact Linux arm64 CUDA (DGX Spark)", | |
| "status": "cross-compiled; NVIDIA runtime conformance pending" | |
| } | |
| ] | |
| }, | |
| "integrations": [ | |
| "integrations/btl3-native", | |
| "integrations/ollama" | |
| ], | |
| "documentation": [ | |
| "docs/launch-btl3-compact.md", | |
| "docs/launch-btl3-cuda.md", | |
| "docs/patched-ollama-and-lmstudio.md", | |
| "docs/btl-3-compact-gguf-exporter.md" | |
| ], | |
| "evidence": [ | |
| "evidence/BTL-3-Compact-AVQ2.report.json", | |
| "evidence/native-load-report.json", | |
| "evidence/rtx-pro-6000-speed-2026-07-19.md", | |
| "evidence/compact-validation.md", | |
| "evidence/thinking-escape-v3-smoke.json", | |
| "evidence/thinking-escape-v3-tool-gate.json", | |
| "evidence/thinking-escape-v3-native-runtime.json", | |
| "evidence/thinking-escape-v3-training-manifest.json" | |
| ], | |
| "tools": [ | |
| "tools/install_consumer_bundle.py" | |
| ], | |
| "licenses": [ | |
| "licenses/LICENSE", | |
| "licenses/LICENSE.runtime", | |
| "licenses/THIRD_PARTY_NOTICES.md" | |
| ], | |
| "stock_ollama_compatible": false, | |
| "stock_lm_studio_engine_compatible": false, | |
| "default_thinking": false, | |
| "canonical_refresh": { | |
| "date": "2026-07-24", | |
| "behavior_adapter": { | |
| "identity": "thinking-escape-r8-step160-v3", | |
| "rank": 8, | |
| "alpha": 16, | |
| "tensor_count": 132, | |
| "payload_bytes": 32440320, | |
| "sha256": "7ba9f80be97287775575022e84e735ac3c38b343c0bb83c5c3b0a86c53d6c881" | |
| }, | |
| "non_behavior_tensors_verified_unchanged": 2284, | |
| "packed_decoder_unchanged": true, | |
| "rank32_output_head_correction_unchanged": true, | |
| "tool_retention": "62/63", | |
| "thinking_final_answer_rate": "5/12", | |
| "thinking_coding_termination": "0/3", | |
| "approved_as_full_thinking_fix": false | |
| } | |
| } | |