tabby-tavern-stack / readme.md
jpanasuk's picture
Remove MCP server (never got working), add SearXNG Redis, update README/DEVLOG, clean start-stack.sh
4ef91a7
|
Raw
History Blame Contribute Delete
9.05 kB
# Tabby-Tavern AI Stack
[![Status](https://img.shields.io/badge/Stack-EXL3%20Ready-success)]() [![Engine](https://img.shields.io/badge/TabbyAPI-ExLlamaV3-blue)]() [![License](https://img.shields.io/badge/License-MIT-green)]()
Welcome to the **Tabby-Tavern** repository! This is a complete, containerized production-grade AI infrastructure stack optimized for high-performance local language models. It natively integrates **TabbyAPI (EXL3 support)**, **SillyTavern**, **Open WebUI**, **Ollama**, and **SearXNG** into a single master orchestration blueprint.
---
## System Requirements & Prerequisites
Before deploying the stack, ensure your host environment meets the following requirements:
* **Operating System:** Linux (Ubuntu/Debian recommended) or Windows via WSL2.
* **GPU:** NVIDIA GPU with sufficient VRAM (e.g., RTX 4070 or higher).
* **Software Stack:**
* Docker & Docker Compose (with NVIDIA Container Toolkit installed for GPU passthrough).
* Python 3.10+ (for utility scripts and Hugging Face CLI tools).
---
## Complete Setup & Installation Guide
Follow these steps to clone, configure, and launch the entire stack from scratch:
### 1. Clone the Repository
Clone your repository directly onto your local machine:
```bash
git clone https://huggingface.co/jpanasuk/tabby-tavern-stack tabby-tavern
cd tabby-tavern
```
### 2. Configure Environment & Security Switches
Ensure your configuration files (located in the `./config/` and `./tabby_config/` directories) are properly mapped.
**Important -- Before First Launch:** You MUST edit the following files to set your own secrets and keys. See the **Security Configuration** section below.
### 3. Add Your EXL3 Models
Place your high-performance ExLlamaV3 (EXL3) model folders directly inside the local model directory:
```bash
# Example path structure
./tabby_models/your-exl3-model-name/
```
### 4. Launch the Master Compose Stack
Spin up all microservices in detached mode using Docker Compose:
```bash
docker compose up -d
```
### 5. Verify Container Health & Logs
Check that all containers (TabbyAPI, SillyTavern, Open WebUI, Ollama, SearXNG) are running smoothly:
```bash
docker compose ps
docker compose logs -f tabbyapi
```
---
## Accessing Your Interfaces
Once the containers are running, you can access your services locally via your browser:
| Service | URL |
|---------|-----|
| SillyTavern Frontend | http://localhost:8000 |
| Open WebUI | http://localhost:3000 |
| TabbyAPI Endpoints | http://localhost:5000 |
| SearXNG Search Engine | http://localhost:8080 |
| Ollama API | http://localhost:11434 |
---
## Repository Architecture Layout
```
tabby-tavern/
β”œβ”€β”€ docker-compose.yml # Master orchestration blueprint (all services + GPU flags)
β”œβ”€β”€ Dockerfile # Custom TabbyAPI image (python3-dev for Triton/CUDA builds)
β”œβ”€β”€ Dockerfile.tabby # Alternate TabbyAPI Dockerfile
β”œβ”€β”€ config.yml # Root-level TabbyAPI config (model, network, security)
β”œβ”€β”€ tabby_config/ # TabbyAPI runtime config (config.yml, api_tokens.yml)
β”œβ”€β”€ tabby_models/ # Storage directory for EXL3/EXL2 model weights
β”œβ”€β”€ sillytavern_config/ # SillyTavern server configuration (config.yaml)
β”œβ”€β”€ sillytavern_data/ # SillyTavern user data (characters, chats, settings)
β”œβ”€β”€ ollama_data/ # Ollama model storage
β”œβ”€β”€ openwebui_data/ # Open WebUI persistent data
β”œβ”€β”€ searxng/ # SearXNG settings
β”œβ”€β”€ searxng_config/ # SearXNG alternate config
β”œβ”€β”€ tabbyAPI-src/ # TabbyAPI source code (for custom builds)
β”œβ”€β”€ Clean_Sampler.json # Recommended sampler preset
β”œβ”€β”€ load-model.sh # Helper script to switch models
β”œβ”€β”€ start-stack.sh # Helper script to start the full stack
└── DEVLOG.md # Development log & engineering history
```
---
## In-Container GPU Tuning
The docker-compose.yml and Dockerfile include CUDA environment variables for maximum inference performance inside containers. The NVIDIA Container Toolkit mounts the host driver -- these variables tune how the container uses it.
### TabbyAPI (EXL3) Container
| Variable | Value | Effect |
|----------|-------|--------|
| `CUDA_VISIBLE_DEVICES` | `0` | Pin to GPU 0 |
| `CUDA_DEVICE_ORDER` | `PCI_BUS_ID` | Consistent device ordering |
| `PYTORCH_CUDA_ALLOC_CONF` | `max_split_size_mb:128` | Reduce memory fragmentation |
| `OMP_NUM_THREADS` | `8` | Match CPU cores for thread pinning |
| `EXLLAMA_GPU_LAYERS` | `999` | Full CUDA layer offloading |
| `EXLLAMA_KV_CACHE` | `q8_0` | 8-bit KV cache quantization |
| `EXLLAMA_FLASH_ATTENTION` | `1` | Enable Flash Attention |
### Ollama Container
| Variable | Value | Effect |
|----------|-------|--------|
| `OLLAMA_FLASH_ATTENTION` | `1` | Enable Flash Attention |
| `OLLAMA_KV_CACHE_TYPE` | `q8_0` | 8-bit KV cache quantization |
| `OLLAMA_NUM_PARALLEL` | `1` | Single-request focus for max throughput |
| `OMP_NUM_THREADS` | `8` | Thread pinning |
### Shared Memory
Both GPU containers use `shm_size: 16g` for large model transfer buffers.
### Host-Level Tuning (not inside Docker)
These must be set on the host before starting containers:
```bash
# GPU power locking (persistent mode)
sudo nvidia-smi -pm 1
# WSL2 kernel tuning (already in /etc/sysctl.d/)
sysctl vm.swappiness=10
echo always > /sys/kernel/mm/transparent_hugepage/enabled
```
---
## Security Configuration (IMPORTANT -- Read Before First Launch)
Several configuration files contain placeholder values that you MUST replace with your own secrets before deploying. Never use the defaults in production.
### 1. TabbyAPI API Keys (`tabby_config/api_tokens.yml`)
Replace `YOUR_TABBY_API_KEY_HERE` and `YOUR_TABBY_ADMIN_KEY_HERE` with strong random strings.
```bash
# Generate a random key:
openssl rand -hex 16
```
### 2. TabbyAPI Config (`config.yml` and `tabby_config/config.yml`)
Replace the `admin_key` and `api_key` placeholder values with your own keys.
### 3. SearXNG Secret Key (`searxng/settings.yml` and `searxng_config/settings.yml`)
Replace `YOUR_SEARXNG_SECRET_KEY_HERE` with a strong random string.
```bash
openssl rand -hex 16
```
### 4. SillyTavern Basic Auth (`sillytavern_config/config.yaml`)
The `basicAuthUser` section has `username` and `password` fields. Change these from the defaults.
### 5. Ollama SSH Keys (`ollama_data/`)
If you previously had SSH keys in `ollama_data/`, they have been excluded from this repo. Generate fresh ones if needed.
---
## Tabby-Tavern Development Log & Engineering History
This devlog tracks the architectural decisions, troubleshooting steps, and configuration milestones achieved while building, debugging, and stabilizing the Tabby-Tavern infrastructure stack.
### Week 1: Core Integration, Engine Iterations & Production Hardening
**1. Hardware & Environment Baseline**
- GPU: NVIDIA GeForce RTX 4070 (VRAM optimization for high-speed token generation).
- Environment: Containerized Linux stack orchestrated via Docker Compose (`docker-compose.yml`) with active CUDA device reservation flags (`capabilities: [gpu]`).
- Primary Models: High-performance EXL3 / ExLlamaV3 formats running on optimized backend runners.
**2. Backend Engine & Quantization Evolution (EXL3 Adoption)**
- The Transition: Upgraded upstream TabbyAPI containers to natively leverage ExLlamaV3 (EXL3), providing optimized memory mapping and maximum token throughput.
- The Resolution: Standardized model pipelines on EXL3 repositories, ensuring instant VRAM loading and seamless API integration with SillyTavern and Open WebUI.
**3. Master Compose Orchestration (docker-compose.yml)**
- Consolidated all microservices (TabbyAPI, SillyTavern, Open WebUI, Ollama, SearXNG, and Redis) into a single, unified `docker-compose.yml` master blueprint.
- Configured local persistent volume mounts for configuration files, model directories, and workspace data containers.
**4. Critical Bugfix: Resolving the TabbyAPI Whitelist Error**
- The Problem: During initial API validation and frontend connection tests, TabbyAPI threw strict whitelist/authorization blocks, rejecting incoming communication requests from external containers or frontend routes due to default security switch constraints.
- The Investigation: Analyzed upstream configuration parameters, environmental switches, and security requirements governing endpoint access control.
- The Resolution: Successfully bypassed and resolved the authorization blocks by modifying the configuration switches and explicitly defining an admin username and password within the environment/config layers. This granted immediate administrative control, unlocked full API endpoint access, and restored seamless pipeline communication across the stack.
**5. Hugging Face Integration**
- Packaged, version-controlled, and pushed the entire production stack (including complete documentation, scripts, and configurations) to remote repositories on Hugging Face.
---
License: MIT