# Tabby-Tavern AI Stack [![Status](https://img.shields.io/badge/Stack-EXL3%20Ready-success)]() [![Engine](https://img.shields.io/badge/TabbyAPI-ExLlamaV3-blue)]() [![License](https://img.shields.io/badge/License-MIT-green)]() Welcome to the **Tabby-Tavern** repository! This is a complete, containerized production-grade AI infrastructure stack optimized for high-performance local language models. It natively integrates **TabbyAPI (EXL3 support)**, **SillyTavern**, **Open WebUI**, **Ollama**, and **SearXNG** into a single master orchestration blueprint. --- ## System Requirements & Prerequisites Before deploying the stack, ensure your host environment meets the following requirements: * **Operating System:** Linux (Ubuntu/Debian recommended) or Windows via WSL2. * **GPU:** NVIDIA GPU with sufficient VRAM (e.g., RTX 4070 or higher). * **Software Stack:** * Docker & Docker Compose (with NVIDIA Container Toolkit installed for GPU passthrough). * Python 3.10+ (for utility scripts and Hugging Face CLI tools). --- ## Complete Setup & Installation Guide Follow these steps to clone, configure, and launch the entire stack from scratch: ### 1. Clone the Repository Clone your repository directly onto your local machine: ```bash git clone https://huggingface.co/jpanasuk/tabby-tavern-stack tabby-tavern cd tabby-tavern ``` ### 2. Configure Environment & Security Switches Ensure your configuration files (located in the `./config/` and `./tabby_config/` directories) are properly mapped. **Important -- Before First Launch:** You MUST edit the following files to set your own secrets and keys. See the **Security Configuration** section below. ### 3. Add Your EXL3 Models Place your high-performance ExLlamaV3 (EXL3) model folders directly inside the local model directory: ```bash # Example path structure ./tabby_models/your-exl3-model-name/ ``` ### 4. Launch the Master Compose Stack Spin up all microservices in detached mode using Docker Compose: ```bash docker compose up -d ``` ### 5. Verify Container Health & Logs Check that all containers (TabbyAPI, SillyTavern, Open WebUI, Ollama, SearXNG) are running smoothly: ```bash docker compose ps docker compose logs -f tabbyapi ``` --- ## Accessing Your Interfaces Once the containers are running, you can access your services locally via your browser: | Service | URL | |---------|-----| | SillyTavern Frontend | http://localhost:8000 | | Open WebUI | http://localhost:3000 | | TabbyAPI Endpoints | http://localhost:5000 | | SearXNG Search Engine | http://localhost:8080 | | Ollama API | http://localhost:11434 | --- ## Repository Architecture Layout ``` tabby-tavern/ ├── docker-compose.yml # Master orchestration blueprint (all services + GPU flags) ├── Dockerfile # Custom TabbyAPI image (python3-dev for Triton/CUDA builds) ├── Dockerfile.tabby # Alternate TabbyAPI Dockerfile ├── config.yml # Root-level TabbyAPI config (model, network, security) ├── tabby_config/ # TabbyAPI runtime config (config.yml, api_tokens.yml) ├── tabby_models/ # Storage directory for EXL3/EXL2 model weights ├── sillytavern_config/ # SillyTavern server configuration (config.yaml) ├── sillytavern_data/ # SillyTavern user data (characters, chats, settings) ├── ollama_data/ # Ollama model storage ├── openwebui_data/ # Open WebUI persistent data ├── searxng/ # SearXNG settings ├── searxng_config/ # SearXNG alternate config ├── tabbyAPI-src/ # TabbyAPI source code (for custom builds) ├── Clean_Sampler.json # Recommended sampler preset ├── load-model.sh # Helper script to switch models ├── start-stack.sh # Helper script to start the full stack └── DEVLOG.md # Development log & engineering history ``` --- ## In-Container GPU Tuning The docker-compose.yml and Dockerfile include CUDA environment variables for maximum inference performance inside containers. The NVIDIA Container Toolkit mounts the host driver -- these variables tune how the container uses it. ### TabbyAPI (EXL3) Container | Variable | Value | Effect | |----------|-------|--------| | `CUDA_VISIBLE_DEVICES` | `0` | Pin to GPU 0 | | `CUDA_DEVICE_ORDER` | `PCI_BUS_ID` | Consistent device ordering | | `PYTORCH_CUDA_ALLOC_CONF` | `max_split_size_mb:128` | Reduce memory fragmentation | | `OMP_NUM_THREADS` | `8` | Match CPU cores for thread pinning | | `EXLLAMA_GPU_LAYERS` | `999` | Full CUDA layer offloading | | `EXLLAMA_KV_CACHE` | `q8_0` | 8-bit KV cache quantization | | `EXLLAMA_FLASH_ATTENTION` | `1` | Enable Flash Attention | ### Ollama Container | Variable | Value | Effect | |----------|-------|--------| | `OLLAMA_FLASH_ATTENTION` | `1` | Enable Flash Attention | | `OLLAMA_KV_CACHE_TYPE` | `q8_0` | 8-bit KV cache quantization | | `OLLAMA_NUM_PARALLEL` | `1` | Single-request focus for max throughput | | `OMP_NUM_THREADS` | `8` | Thread pinning | ### Shared Memory Both GPU containers use `shm_size: 16g` for large model transfer buffers. ### Host-Level Tuning (not inside Docker) These must be set on the host before starting containers: ```bash # GPU power locking (persistent mode) sudo nvidia-smi -pm 1 # WSL2 kernel tuning (already in /etc/sysctl.d/) sysctl vm.swappiness=10 echo always > /sys/kernel/mm/transparent_hugepage/enabled ``` --- ## Security Configuration (IMPORTANT -- Read Before First Launch) Several configuration files contain placeholder values that you MUST replace with your own secrets before deploying. Never use the defaults in production. ### 1. TabbyAPI API Keys (`tabby_config/api_tokens.yml`) Replace `YOUR_TABBY_API_KEY_HERE` and `YOUR_TABBY_ADMIN_KEY_HERE` with strong random strings. ```bash # Generate a random key: openssl rand -hex 16 ``` ### 2. TabbyAPI Config (`config.yml` and `tabby_config/config.yml`) Replace the `admin_key` and `api_key` placeholder values with your own keys. ### 3. SearXNG Secret Key (`searxng/settings.yml` and `searxng_config/settings.yml`) Replace `YOUR_SEARXNG_SECRET_KEY_HERE` with a strong random string. ```bash openssl rand -hex 16 ``` ### 4. SillyTavern Basic Auth (`sillytavern_config/config.yaml`) The `basicAuthUser` section has `username` and `password` fields. Change these from the defaults. ### 5. Ollama SSH Keys (`ollama_data/`) If you previously had SSH keys in `ollama_data/`, they have been excluded from this repo. Generate fresh ones if needed. --- ## Tabby-Tavern Development Log & Engineering History This devlog tracks the architectural decisions, troubleshooting steps, and configuration milestones achieved while building, debugging, and stabilizing the Tabby-Tavern infrastructure stack. ### Week 1: Core Integration, Engine Iterations & Production Hardening **1. Hardware & Environment Baseline** - GPU: NVIDIA GeForce RTX 4070 (VRAM optimization for high-speed token generation). - Environment: Containerized Linux stack orchestrated via Docker Compose (`docker-compose.yml`) with active CUDA device reservation flags (`capabilities: [gpu]`). - Primary Models: High-performance EXL3 / ExLlamaV3 formats running on optimized backend runners. **2. Backend Engine & Quantization Evolution (EXL3 Adoption)** - The Transition: Upgraded upstream TabbyAPI containers to natively leverage ExLlamaV3 (EXL3), providing optimized memory mapping and maximum token throughput. - The Resolution: Standardized model pipelines on EXL3 repositories, ensuring instant VRAM loading and seamless API integration with SillyTavern and Open WebUI. **3. Master Compose Orchestration (docker-compose.yml)** - Consolidated all microservices (TabbyAPI, SillyTavern, Open WebUI, Ollama, SearXNG, and Redis) into a single, unified `docker-compose.yml` master blueprint. - Configured local persistent volume mounts for configuration files, model directories, and workspace data containers. **4. Critical Bugfix: Resolving the TabbyAPI Whitelist Error** - The Problem: During initial API validation and frontend connection tests, TabbyAPI threw strict whitelist/authorization blocks, rejecting incoming communication requests from external containers or frontend routes due to default security switch constraints. - The Investigation: Analyzed upstream configuration parameters, environmental switches, and security requirements governing endpoint access control. - The Resolution: Successfully bypassed and resolved the authorization blocks by modifying the configuration switches and explicitly defining an admin username and password within the environment/config layers. This granted immediate administrative control, unlocked full API endpoint access, and restored seamless pipeline communication across the stack. **5. Hugging Face Integration** - Packaged, version-controlled, and pushed the entire production stack (including complete documentation, scripts, and configurations) to remote repositories on Hugging Face. --- License: MIT