Text Generation
GGUF
llama.cpp
agent
coding
reasoning
tool-use
function-calling
quantized
cuda
metal
conversational
Instructions to use badtheorylabs/BTL-3-Compact with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use badtheorylabs/BTL-3-Compact with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./llama-cli -hf badtheorylabs/BTL-3-Compact
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./build/bin/llama-cli -hf badtheorylabs/BTL-3-Compact
Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- LM Studio
- Jan
- vLLM
How to use badtheorylabs/BTL-3-Compact with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "badtheorylabs/BTL-3-Compact" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badtheorylabs/BTL-3-Compact", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Ollama
How to use badtheorylabs/BTL-3-Compact with Ollama:
ollama run hf.co/badtheorylabs/BTL-3-Compact
- Unsloth Studio
How to use badtheorylabs/BTL-3-Compact with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
- Pi
How to use badtheorylabs/BTL-3-Compact with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "badtheorylabs/BTL-3-Compact" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use badtheorylabs/BTL-3-Compact with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default badtheorylabs/BTL-3-Compact
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use badtheorylabs/BTL-3-Compact with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "badtheorylabs/BTL-3-Compact" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use badtheorylabs/BTL-3-Compact with Docker Model Runner:
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Lemonade
How to use badtheorylabs/BTL-3-Compact with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull badtheorylabs/BTL-3-Compact
Run and chat with the model
lemonade run user.BTL-3-Compact-{{QUANT_TAG}}List all available models
lemonade list
| import { spawn, type ChildProcess } from "node:child_process"; | |
| import { homedir } from "node:os"; | |
| import { dirname, join } from "node:path"; | |
| type RunnerOptions = { | |
| baseUrl: string; | |
| runnerPath: string; | |
| modelPath: string; | |
| }; | |
| let child: ChildProcess | undefined; | |
| let startup: Promise<void> | undefined; | |
| let stderrTail = ""; | |
| let childError: Error | undefined; | |
| function installedPaths(): { runner: string; model: string } { | |
| if (process.platform === "win32") { | |
| const root = process.env.LOCALAPPDATA ?? join(homedir(), "AppData", "Local"); | |
| return { | |
| runner: join(root, "BTL3", "libexec", "llama-server.exe"), | |
| model: join(root, "BTL3", "model", "BTL-3-Compact-AVQ2.gguf"), | |
| }; | |
| } | |
| const root = process.env.XDG_DATA_HOME ?? join(homedir(), ".local", "share"); | |
| return { | |
| runner: join(root, "btl3", "libexec", "llama-server"), | |
| model: join(root, "btl3", "model", "BTL-3-Compact-AVQ2.gguf"), | |
| }; | |
| } | |
| function endpoint(baseUrl: string, path: string): URL { | |
| const url = new URL(baseUrl); | |
| url.pathname = path; | |
| url.search = ""; | |
| return url; | |
| } | |
| async function healthy(baseUrl: string): Promise<boolean> { | |
| try { | |
| const response = await fetch(endpoint(baseUrl, "/health"), { | |
| signal: AbortSignal.timeout(1_500), | |
| }); | |
| return response.ok; | |
| } catch { | |
| return false; | |
| } | |
| } | |
| function runnerArguments(baseUrl: string, modelPath: string): string[] { | |
| const url = new URL(baseUrl); | |
| const port = url.port || "8080"; | |
| return [ | |
| "--model", modelPath, | |
| "--host", url.hostname, | |
| "--port", port, | |
| "--no-webui", | |
| "--offline", | |
| "-np", "1", | |
| ]; | |
| } | |
| async function waitUntilReady(baseUrl: string): Promise<void> { | |
| const deadline = Date.now() + 60_000; | |
| while (Date.now() < deadline) { | |
| if (await healthy(baseUrl)) return; | |
| if (childError) { | |
| throw new Error(`BTL-3 runner failed to start: ${childError.message}`); | |
| } | |
| if (child?.exitCode !== null && child?.exitCode !== undefined) { | |
| throw new Error(`BTL-3 runner exited (${child.exitCode}): ${stderrTail}`); | |
| } | |
| await new Promise(resolve => setTimeout(resolve, 250)); | |
| } | |
| throw new Error(`BTL-3 runner did not become ready: ${stderrTail}`); | |
| } | |
| async function start(options: RunnerOptions): Promise<void> { | |
| if (await healthy(options.baseUrl)) return; | |
| const defaults = installedPaths(); | |
| const runner = options.runnerPath || process.env.BTL3_RUNNER_PATH || defaults.runner; | |
| const model = options.modelPath || process.env.BTL3_MODEL_PATH || defaults.model; | |
| const root = dirname(dirname(runner)); | |
| const libraryPath = join(root, "lib"); | |
| const env = { ...process.env }; | |
| if (process.platform === "win32") { | |
| env.PATH = `${libraryPath};${env.PATH ?? ""}`; | |
| env.GGML_BACKEND_PATH = join(libraryPath, "ggml-cuda.dll"); | |
| } else { | |
| env.LD_LIBRARY_PATH = | |
| `${libraryPath}${env.LD_LIBRARY_PATH ? `:${env.LD_LIBRARY_PATH}` : ""}`; | |
| env.GGML_BACKEND_PATH = join(libraryPath, "libggml-cuda.so"); | |
| } | |
| stderrTail = ""; | |
| childError = undefined; | |
| child = spawn(runner, runnerArguments(options.baseUrl, model), { | |
| windowsHide: true, | |
| stdio: ["ignore", "ignore", "pipe"], | |
| env, | |
| }); | |
| child.stderr?.on("data", data => { | |
| stderrTail = (stderrTail + String(data)).slice(-4_096); | |
| }); | |
| child.once("error", error => { | |
| childError = error; | |
| stderrTail = (stderrTail + error.message).slice(-4_096); | |
| }); | |
| await waitUntilReady(options.baseUrl); | |
| } | |
| export async function ensureNativeRunner(options: RunnerOptions): Promise<void> { | |
| if (await healthy(options.baseUrl)) return; | |
| startup ??= start(options).finally(() => { | |
| startup = undefined; | |
| }); | |
| await startup; | |
| } | |
| process.once("exit", () => { | |
| child?.kill(); | |
| }); | |