Text Generation
GGUF
llama.cpp
agent
coding
reasoning
tool-use
function-calling
quantized
cuda
metal
conversational
Instructions to use badtheorylabs/BTL-3-Compact with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use badtheorylabs/BTL-3-Compact with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: llama cli -hf badtheorylabs/BTL-3-Compact
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./llama-cli -hf badtheorylabs/BTL-3-Compact
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf badtheorylabs/BTL-3-Compact # Run inference directly in the terminal: ./build/bin/llama-cli -hf badtheorylabs/BTL-3-Compact
Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- LM Studio
- Jan
- vLLM
How to use badtheorylabs/BTL-3-Compact with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "badtheorylabs/BTL-3-Compact" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "badtheorylabs/BTL-3-Compact", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Ollama
How to use badtheorylabs/BTL-3-Compact with Ollama:
ollama run hf.co/badtheorylabs/BTL-3-Compact
- Unsloth Studio
How to use badtheorylabs/BTL-3-Compact with Unsloth Studio:
Install Unsloth Studio (macOS, Linux, WSL)
curl -fsSL https://unsloth.ai/install.sh | sh # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Install Unsloth Studio (Windows)
irm https://unsloth.ai/install.ps1 | iex # Run unsloth studio unsloth studio -H 0.0.0.0 -p 8888 # Then open http://localhost:8888 in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
Using HuggingFace Spaces for Unsloth
# No setup required # Open https://huggingface.co/spaces/unsloth/studio in your browser # Search for badtheorylabs/BTL-3-Compact to start chatting
- Pi
How to use badtheorylabs/BTL-3-Compact with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure the model in Pi
# Install Pi: npm install -g @mariozechner/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "badtheorylabs/BTL-3-Compact" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Hermes Agent new
How to use badtheorylabs/BTL-3-Compact with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default badtheorylabs/BTL-3-Compact
Run Hermes
hermes
- Atomic Chat new
- OpenClaw new
How to use badtheorylabs/BTL-3-Compact with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf badtheorylabs/BTL-3-Compact
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "badtheorylabs/BTL-3-Compact" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
- Docker Model Runner
How to use badtheorylabs/BTL-3-Compact with Docker Model Runner:
docker model run hf.co/badtheorylabs/BTL-3-Compact
- Lemonade
How to use badtheorylabs/BTL-3-Compact with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull badtheorylabs/BTL-3-Compact
Run and chat with the model
lemonade run user.BTL-3-Compact-{{QUANT_TAG}}List all available models
lemonade list
| import { | |
| type Chat, | |
| type GeneratorController, | |
| type InferParsedConfig, | |
| type PluginContext, | |
| } from "@lmstudio/sdk"; | |
| import OpenAI from "openai"; | |
| import type { | |
| ChatCompletionChunk, | |
| ChatCompletionMessageParam, | |
| ChatCompletionMessageToolCall, | |
| ChatCompletionTool, | |
| } from "openai/resources/chat/completions"; | |
| import { globalConfigSchematics } from "./config"; | |
| import { ensureNativeRunner } from "./nativeRunner"; | |
| type Config = InferParsedConfig<typeof globalConfigSchematics>; | |
| type ToolState = { | |
| id: string; | |
| name: string; | |
| argumentFragments: string[]; | |
| }; | |
| function messages(history: Chat): ChatCompletionMessageParam[] { | |
| const result: ChatCompletionMessageParam[] = []; | |
| for (const message of history) { | |
| switch (message.getRole()) { | |
| case "system": | |
| result.push({ role: "system", content: message.getText() }); | |
| break; | |
| case "user": | |
| result.push({ role: "user", content: message.getText() }); | |
| break; | |
| case "assistant": { | |
| const toolCalls: ChatCompletionMessageToolCall[] = message | |
| .getToolCallRequests() | |
| .map(call => ({ | |
| id: call.id ?? "", | |
| type: "function", | |
| function: { | |
| name: call.name, | |
| arguments: JSON.stringify(call.arguments ?? {}), | |
| }, | |
| })); | |
| result.push({ | |
| role: "assistant", | |
| content: message.getText(), | |
| ...(toolCalls.length > 0 ? { tool_calls: toolCalls } : {}), | |
| }); | |
| break; | |
| } | |
| case "tool": | |
| for (const toolResult of message.getToolCallResults()) { | |
| result.push({ | |
| role: "tool", | |
| tool_call_id: toolResult.toolCallId ?? "", | |
| content: toolResult.content, | |
| }); | |
| } | |
| break; | |
| } | |
| } | |
| return result; | |
| } | |
| function tools(ctl: GeneratorController): ChatCompletionTool[] | undefined { | |
| const result = ctl.getToolDefinitions().map<ChatCompletionTool>(tool => ({ | |
| type: "function", | |
| function: { | |
| name: tool.function.name, | |
| description: tool.function.description, | |
| parameters: tool.function.parameters ?? {}, | |
| }, | |
| })); | |
| return result.length > 0 ? result : undefined; | |
| } | |
| function asError(error: unknown): Error { | |
| return error instanceof Error ? error : new Error(String(error)); | |
| } | |
| function finishTool(ctl: GeneratorController, state: ToolState): void { | |
| ctl.toolCallGenerationStarted({ toolCallId: state.id }); | |
| if (state.name) { | |
| ctl.toolCallGenerationNameReceived(state.name); | |
| } | |
| for (const fragment of state.argumentFragments) { | |
| ctl.toolCallGenerationArgumentFragmentGenerated(fragment); | |
| } | |
| try { | |
| const parsed = JSON.parse( | |
| state.argumentFragments.join("") || "{}", | |
| ) as Record<string, unknown>; | |
| ctl.toolCallGenerationEnded({ | |
| type: "function", | |
| id: state.id, | |
| name: state.name, | |
| arguments: parsed, | |
| }); | |
| } catch (error) { | |
| ctl.toolCallGenerationFailed(asError(error)); | |
| } | |
| } | |
| function bufferToolDelta( | |
| pending: Map<number, ToolState>, | |
| call: ChatCompletionChunk.Choice.Delta.ToolCall, | |
| ): void { | |
| const state = pending.get(call.index) ?? { | |
| id: call.id ?? `call_${call.index}`, | |
| name: "", | |
| argumentFragments: [], | |
| }; | |
| if (call.id) state.id = call.id; | |
| if (call.function?.name) { | |
| state.name += call.function.name; | |
| } | |
| if (call.function?.arguments) { | |
| state.argumentFragments.push(call.function.arguments); | |
| } | |
| pending.set(call.index, state); | |
| } | |
| async function generate( | |
| ctl: GeneratorController, | |
| history: Chat, | |
| config: Config, | |
| ): Promise<void> { | |
| if (config.get("autoStart")) { | |
| await ensureNativeRunner({ | |
| baseUrl: config.get("baseUrl"), | |
| runnerPath: config.get("runnerPath"), | |
| modelPath: config.get("modelPath"), | |
| }); | |
| } | |
| const client = new OpenAI({ | |
| apiKey: config.get("apiKey") || "btl3-local", | |
| baseURL: config.get("baseUrl"), | |
| }); | |
| const toolDefinitions = tools(ctl); | |
| const pending = new Map<number, ToolState>(); | |
| try { | |
| ctl.abortSignal.throwIfAborted(); | |
| const stream = await client.chat.completions.create( | |
| { | |
| model: "BTL-3", | |
| messages: messages(history), | |
| tools: toolDefinitions, | |
| ...(toolDefinitions ? { parallel_tool_calls: true } : {}), | |
| stream: true, | |
| }, | |
| { signal: ctl.abortSignal }, | |
| ); | |
| for await (const chunk of stream) { | |
| ctl.abortSignal.throwIfAborted(); | |
| const delta = chunk.choices[0]?.delta; | |
| if (!delta) continue; | |
| const reasoning = ( | |
| delta as typeof delta & { reasoning_content?: string } | |
| ).reasoning_content; | |
| if (reasoning) { | |
| ctl.fragmentGenerated(reasoning, { reasoningType: "reasoning" }); | |
| } | |
| if (delta.content) { | |
| ctl.fragmentGenerated(delta.content); | |
| } | |
| for (const call of delta.tool_calls ?? []) { | |
| bufferToolDelta(pending, call); | |
| } | |
| } | |
| for (const [, state] of [...pending].sort(([a], [b]) => a - b)) { | |
| finishTool(ctl, state); | |
| } | |
| } catch (error) { | |
| throw asError(error); | |
| } | |
| } | |
| export async function main(context: PluginContext): Promise<void> { | |
| context.withGlobalConfigSchematics(globalConfigSchematics); | |
| context.withGenerator(async (ctl, history) => { | |
| const config = ctl.getGlobalPluginConfig(globalConfigSchematics); | |
| await generate(ctl, history, config); | |
| }); | |
| } | |