Text Generation
GGUF
llama.cpp
agent
coding
reasoning
tool-use
function-calling
quantized
cuda
metal
conversational
Instructions to use badtheorylabs/BTL-3-Compact with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use badtheorylabs/BTL-3-Compact with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./llama-cli -hf badtheorylabs/BTL-3-Compact
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./build/bin/llama-cli -hf badtheorylabs/BTL-3-Compact
Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- LM Studio
- Jan
- vLLM
How to use badtheorylabs/BTL-3-Compact with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "badtheorylabs/BTL-3-Compact" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badtheorylabs/BTL-3-Compact", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Ollama
How to use badtheorylabs/BTL-3-Compact with Ollama:
ollama run hf.co/badtheorylabs/BTL-3-Compact
- Unsloth Studio
How to use badtheorylabs/BTL-3-Compact with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
- Pi
How to use badtheorylabs/BTL-3-Compact with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "badtheorylabs/BTL-3-Compact" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use badtheorylabs/BTL-3-Compact with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default badtheorylabs/BTL-3-Compact
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use badtheorylabs/BTL-3-Compact with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "badtheorylabs/BTL-3-Compact" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use badtheorylabs/BTL-3-Compact with Docker Model Runner:
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Lemonade
How to use badtheorylabs/BTL-3-Compact with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull badtheorylabs/BTL-3-Compact
Run and chat with the model
lemonade run user.BTL-3-Compact-{{QUANT_TAG}}List all available models
lemonade list
File size: 3,692 Bytes
274ffba | 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57 58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102 103 104 105 106 107 108 109 110 111 112 113 114 115 116 117 118 119 120 121 | import { spawn, type ChildProcess } from "node:child_process";
import { homedir } from "node:os";
import { dirname, join } from "node:path";
type RunnerOptions = {
baseUrl: string;
runnerPath: string;
modelPath: string;
};
let child: ChildProcess | undefined;
let startup: Promise<void> | undefined;
let stderrTail = "";
let childError: Error | undefined;
function installedPaths(): { runner: string; model: string } {
if (process.platform === "win32") {
const root = process.env.LOCALAPPDATA ?? join(homedir(), "AppData", "Local");
return {
runner: join(root, "BTL3", "libexec", "llama-server.exe"),
model: join(root, "BTL3", "model", "BTL-3-Compact-AVQ2.gguf"),
};
}
const root = process.env.XDG_DATA_HOME ?? join(homedir(), ".local", "share");
return {
runner: join(root, "btl3", "libexec", "llama-server"),
model: join(root, "btl3", "model", "BTL-3-Compact-AVQ2.gguf"),
};
}
function endpoint(baseUrl: string, path: string): URL {
const url = new URL(baseUrl);
url.pathname = path;
url.search = "";
return url;
}
async function healthy(baseUrl: string): Promise<boolean> {
try {
const response = await fetch(endpoint(baseUrl, "/health"), {
signal: AbortSignal.timeout(1_500),
});
return response.ok;
} catch {
return false;
}
}
function runnerArguments(baseUrl: string, modelPath: string): string[] {
const url = new URL(baseUrl);
const port = url.port || "8080";
return [
"--model", modelPath,
"--host", url.hostname,
"--port", port,
"--no-webui",
"--offline",
"-np", "1",
];
}
async function waitUntilReady(baseUrl: string): Promise<void> {
const deadline = Date.now() + 60_000;
while (Date.now() < deadline) {
if (await healthy(baseUrl)) return;
if (childError) {
throw new Error(`BTL-3 runner failed to start: ${childError.message}`);
}
if (child?.exitCode !== null && child?.exitCode !== undefined) {
throw new Error(`BTL-3 runner exited (${child.exitCode}): ${stderrTail}`);
}
await new Promise(resolve => setTimeout(resolve, 250));
}
throw new Error(`BTL-3 runner did not become ready: ${stderrTail}`);
}
async function start(options: RunnerOptions): Promise<void> {
if (await healthy(options.baseUrl)) return;
const defaults = installedPaths();
const runner = options.runnerPath || process.env.BTL3_RUNNER_PATH || defaults.runner;
const model = options.modelPath || process.env.BTL3_MODEL_PATH || defaults.model;
const root = dirname(dirname(runner));
const libraryPath = join(root, "lib");
const env = { ...process.env };
if (process.platform === "win32") {
env.PATH = `${libraryPath};${env.PATH ?? ""}`;
env.GGML_BACKEND_PATH = join(libraryPath, "ggml-cuda.dll");
} else {
env.LD_LIBRARY_PATH =
`${libraryPath}${env.LD_LIBRARY_PATH ? `:${env.LD_LIBRARY_PATH}` : ""}`;
env.GGML_BACKEND_PATH = join(libraryPath, "libggml-cuda.so");
}
stderrTail = "";
childError = undefined;
child = spawn(runner, runnerArguments(options.baseUrl, model), {
windowsHide: true,
stdio: ["ignore", "ignore", "pipe"],
env,
});
child.stderr?.on("data", data => {
stderrTail = (stderrTail + String(data)).slice(-4_096);
});
child.once("error", error => {
childError = error;
stderrTail = (stderrTail + error.message).slice(-4_096);
});
await waitUntilReady(options.baseUrl);
}
export async function ensureNativeRunner(options: RunnerOptions): Promise<void> {
if (await healthy(options.baseUrl)) return;
startup ??= start(options).finally(() => {
startup = undefined;
});
await startup;
}
process.once("exit", () => {
child?.kill();
});
|